Wikitech
labswiki
https://wikitech.wikimedia.org/wiki/Main_Page
MediaWiki 1.47.0-wmf.21
first-letter
Media
Special
Talk
User
User talk
Wikitech
Wikitech talk
File
File talk
MediaWiki
MediaWiki talk
Template
Template talk
Help
Help talk
Category
Category talk
Obsolete
Obsolete talk
OfficeIT
OfficeIT talk
Tool
Tool talk
Nova Resource
Nova Resource Talk
Heira
Heira Talk
TimedText
TimedText talk
Module
Module talk
Server Admin Log
0
7919
2461120
2461119
2026-09-26T16:29:50Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
2461120
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
kmuktijg2axus0qbib1v4rehmjlusw1
2461121
2461120
2026-09-26T16:30:15Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
2461121
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
68fdiqjom1nj2b7xxauc5csc63mghgf
2461122
2461121
2026-09-26T16:30:16Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
2461122
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
axa9nmba15o7s2isif3hw4ht9hfle51
2461123
2461122
2026-09-26T16:30:46Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
2461123
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
figzlg63o1gatvpurqdyoc7apqyjbsq
2461124
2461123
2026-09-26T16:37:19Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
2461124
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
5djwsspjvson3ijdvuimzta7mshjdew
2461125
2461124
2026-09-26T16:37:21Z
Stashbot
7414
ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
2461125
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
encxgm1qtdg7wr3fff931zqcakfwyw0
2461126
2461125
2026-09-26T16:37:22Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
2461126
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
ko9hct8onfvqwcihwwenvy0s0mvndul
2461127
2461126
2026-09-26T16:37:25Z
Stashbot
7414
ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
2461127
wikitext
text/x-wiki
== 2026-09-26 ==
* 16:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
d0tudy12ti9vsz37jrqzaqibwtrsb5f
2461133
2461127
2026-09-26T21:26:09Z
Stashbot
7414
krinkle@deploy1003: Started deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296
2461133
wikitext
text/x-wiki
== 2026-09-26 ==
* 21:26 krinkle@deploy1003: Started deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296
* 16:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
swzfaa8drjbr55gef334vcyg0fw961x
2461134
2461133
2026-09-26T21:26:18Z
Stashbot
7414
krinkle@deploy1003: Finished deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296 (duration: 00m 09s)
2461134
wikitext
text/x-wiki
== 2026-09-26 ==
* 21:26 krinkle@deploy1003: Finished deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296 (duration: 00m 09s)
* 21:26 krinkle@deploy1003: Started deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296
* 16:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
kwb61d7bwkcwpt25pznepj0und53hru
2461137
2461134
2026-09-27T02:00:22Z
Stashbot
7414
mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2461137
wikitext
text/x-wiki
== 2026-09-27 ==
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-26 ==
* 21:26 krinkle@deploy1003: Finished deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296 (duration: 00m 09s)
* 21:26 krinkle@deploy1003: Started deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296
* 16:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
hk5js5v7gm1a6y3neah8e6mrosumax6
2461138
2461137
2026-09-27T02:07:55Z
Stashbot
7414
mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 32s)
2461138
wikitext
text/x-wiki
== 2026-09-27 ==
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 32s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-26 ==
* 21:26 krinkle@deploy1003: Finished deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296 (duration: 00m 09s)
* 21:26 krinkle@deploy1003: Started deploy [performance/arc-lamp@68349ee]: https://gerrit.wikimedia.org/r/c/performance/arc-lamp/+/1345296
* 16:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:37 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:29 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:07 oblivian@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] (duration: 10m 53s)
* 08:02 oblivian@deploy1003: oblivian: Continuing with deployment
* 08:00 oblivian@deploy1003: oblivian: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:56 oblivian@deploy1003: Started scap sync-world: Backport for [[gerrit:1345258{{!}}Revert "Disable Score exec"]]
* 07:52 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 07:50 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 07:46 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 07:44 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 07:42 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 06:30 oblivian@deploy1003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 06:29 oblivian@deploy1003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 06:28 oblivian@deploy1003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 06:27 oblivian@deploy1003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 03:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] (duration: 11m 01s)
* 03:31 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 03:30 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 03:26 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1345252{{!}}Disable Score exec (T439297 T438443)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 13s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-25 ==
* 23:15 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 22:51 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:51 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:47 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:46 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:39 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host cloudcephosd1055.eqiad.wmnet with OS bookworm
* 18:27 krinkle@deploy1003: Finished deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values (duration: 00m 11s)
* 18:27 krinkle@deploy1003: Started deploy [statsv/statsv@df3ebff]: [[phab:T439183|T439183]]: Accept dot, plus, hyphen in label values
* 17:59 cdanis@dns1004: END - running authdns-update
* 17:57 cdanis@dns1004: START - running authdns-update
* 15:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2001.*
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:03 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:02 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 15:01 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:57 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2001.codfw.wmnet with OS trixie
* 14:56 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:56 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1017.eqiad.wmnet
* 14:51 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:51 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1040.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1041.eqiad.wmnet
* 14:49 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet
* 14:46 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1040.eqiad.wmnet
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:44 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:43 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1041.eqiad.wmnet
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:41 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 14:38 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:33 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2001.codfw.wmnet with reason: host reimage
* 14:32 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 14:30 dkertesz: moved haproxy stat file from /var/lib/haproxy/stats-file to /run/haproxy/ in cp7001,cp7011 - [[phab:T343000|T343000]]
* 14:29 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 14:23 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,name=cp2059.*
* 14:18 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:17 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 14:15 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 14:14 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2001.codfw.wmnet with OS trixie
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:13 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 14:06 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 14:06 brouberol@cumin1004: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:05 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1041.eqiad.wmnet with reason: host reimage
* 14:02 brouberol@cumin1004: conftool action : set/weight=10; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:01 brouberol@cumin1004: conftool action : set/pooled=yes; selector: name=dse-k8s-worker1039.eqiad.wmnet
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-main,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-logging-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics-external,name=codfw
* 14:00 atsuko@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=eventgate-analytics,name=codfw
* 14:00 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:59 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1040.eqiad.wmnet with reason: host reimage
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:58 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:55 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 13:54 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1041.eqiad.wmnet with OS bookworm
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:53 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1041
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1041 on all recursors
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:52 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:52 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 13:51 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1003 to dse-k8s-worker1041 - brouberol@cumin1004"
* 13:51 zabe: clone wbc_entity_usage from local cluster to x1 for all wikidata client wikis # [[phab:T438750|T438750]]
* 13:50 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 13:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 13:47 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:47 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1003 to dse-k8s-worker1041
* 13:46 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:46 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1019.*<nowiki>}</nowiki> and A:lvs
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1039.eqiad.wmnet
* 13:45 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1040.eqiad.wmnet with OS bookworm
* 13:44 vgutierrez@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:44 vgutierrez@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs1020.*<nowiki>}</nowiki> and A:lvs
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1040
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1040 on all recursors
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1002 to dse-k8s-worker1040 - brouberol@cumin1004"
* 13:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 13:34 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1002 to dse-k8s-worker1040
* 13:29 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in codfw: return to active/active
* 13:24 Emperor: repool sessionstore in codfw
* 13:24 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in codfw: return to active/active
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 13:24 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:22 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 13:15 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 13:03 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:59 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp2049.codfw.wmnet
* 12:58 jclark@cumin1004: START - Cookbook sre.hosts.provision for host wikikube-worker1152.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:52 brouberol@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on dse-k8s-worker1039.eqiad.wmnet with reason: host reimage
* 12:47 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 12:47 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 12:45 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:42 brouberol@cumin1004: START - Cookbook sre.hosts.reimage for host dse-k8s-worker1039.eqiad.wmnet with OS bookworm
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:40 brouberol@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1039
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: START - Cookbook sre.dns.wipe-cache dse-k8s-worker1039 on all recursors
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:38 brouberol@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming ganeti-jumbo1001 to dse-k8s-worker1039 - brouberol@cumin1004"
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cumin1003.eqiad.wmnet
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:34 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:34 brouberol@cumin1004: START - Cookbook sre.dns.netbox
* 12:33 brouberol@cumin1004: START - Cookbook sre.hosts.rename from ganeti-jumbo1001 to dse-k8s-worker1039
* 12:26 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cumin1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 12:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 12:13 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts cumin1003.eqiad.wmnet
* 11:41 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:40 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:revalidateLinkRecommendations.php --olderThan=1790175600 --verbose # [[phab:T438366|T438366]]
* 11:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1002.eqiad.wmnet
* 11:20 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:19 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 11:17 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 10:58 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:53 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 10:52 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 09:49 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 09:48 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 09:43 tappof: reset modified_attributes for hosts and services that fully match the Puppet configuration in Icinga - [[phab:T439105|T439105]]
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 09:36 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 09:35 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts build2001.codfw.wmnet
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:28 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:11 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-worker1207.eqiad.wmnet
* 09:01 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: build2001.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:57 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host an-worker1207.eqiad.wmnet
* 08:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:52 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts build2001.codfw.wmnet
* 08:24 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 08:23 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 08:22 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 08:20 vgutierrez@puppetserver1001: conftool action : set/weight=1; selector: dc=codfw,name=cp2059.*
* 08:15 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,name=cp2059.*
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:58 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 05:21 ryankemper: [Cirrus] Stumble across orphaned index `sawikisource_content_1784136042`, deleted. The real index is `sawikisource_content_1784136826` which I've obviously left untouched
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 38s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 01:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 01:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 01:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 01:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 00:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 00:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 00:14 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:14 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 00:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 00:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
== 2026-09-24 ==
* 23:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 23:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 23:50 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:32 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 23:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:20 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 23:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 22:41 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:41 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 22:28 ryankemper: [WDQS] Expanding match in https://requestctl.wikimedia.org/pattern/ua/rocks to test a likely block candidate
* {{safesubst:SAL entry|1=22:27 egardner@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on t}}
* 22:22 egardner@deploy1003: volker-e, egardner, jdlrobson: Continuing with deployment
* 22:21 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 22:19 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 22:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* {{safesubst:SAL entry|1=22:14 egardner@deploy1003: volker-e, egardner, jdlrobson: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment}}
* {{safesubst:SAL entry|1=22:10 egardner@deploy1003: Started scap sync-world: Backport for [[gerrit:1344049{{!}}ReaderExperiments: Set the preferred-sources debug flag on testwiki (T436692)]], [[gerrit:1344050{{!}}ReaderExperiments: Drop the stale ShareHighlight config var (T424764)]], [[gerrit:1344118{{!}}Enable ReadingList CTA on Minerva for our test wikis (inc beta cluster) (T438779)]], [[gerrit:1343560{{!}}Revert "Enable Reading Recommendations experiment on te}}
* 22:04 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 22:02 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 22:00 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 22:00 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 21:58 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4052.*
* 21:57 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 21:53 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:53 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 21:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 21:46 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:46 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1016.eqiad.wmnet
* 21:45 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] (duration: 17m 05s)
* 21:42 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp4052.ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:40 catrope@deploy1003: catrope: Continuing with deployment
* 21:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1016.eqiad.wmnet
* 21:34 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 21:34 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 21:33 catrope@deploy1003: catrope: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:28 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344795{{!}}Catch newline character in UserMailer to prevent it from allowing bad actors to create an additional header (T434545)]]
* 21:28 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] (duration: 19m 27s)
* 21:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 21:22 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Continuing with deployment
* 21:12 catrope@deploy1003: cjming, mfossati, catrope, mlitn: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]] synced to the testservers (see https://wi
* 21:08 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344406{{!}}ext.wikimediaEvents.testKitchen: Add withContext helper (T438898)]], [[gerrit:1344716{{!}}ReaderExperiments: add dewiki and svwiki (T438072)]], [[gerrit:1344740{{!}}Image Browsing carousel: taps outside the preview dialog should close it (T439006)]], [[gerrit:1344752{{!}}Cap the dialog viewport (T439007)]]
* 21:04 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] (duration: 10m 45s)
* 21:03 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4051.*
* 21:02 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4041.*
* 20:58 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Continuing with deployment
* 20:57 catrope@deploy1003: catrope, ebernhardson, jgiannelos: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:57 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 20:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:53 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344750{{!}}Revert "cirrus: Send more_like traffic to eqiad"]], [[gerrit:1344329{{!}}prv: Enable parsoid rendering for 5 wikis (T438998)]]
* 20:50 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:49 catrope@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] (duration: 10m 19s)
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 20:47 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 20:44 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp405[1-2].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:44 catrope@deploy1003: catrope: Continuing with deployment
* 20:43 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp404[1-6].ulsfo.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 20:43 catrope@deploy1003: catrope: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:39 catrope@deploy1003: Started scap sync-world: Backport for [[gerrit:1344388{{!}}HookHandler: Guard against recovery code expiry being null (T438593)]]
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:25 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 20:19 kemayo@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] (duration: 11m 23s)
* 20:14 kemayo@deploy1003: kemayo: Continuing with deployment
* 20:12 kemayo@deploy1003: kemayo: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 20:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 20:08 kemayo@deploy1003: Started scap sync-world: Backport for [[gerrit:1344714{{!}}EditCheck: add some statsv tracking of check/suggestion actions (T438916)]]
* 20:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:57 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:56 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:50 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 19:46 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 19:36 ryankemper: [Cirrus] All cirrus pools are serving again. Actively monitoring while the system returns to equilibrium, but all initial indications are that things are as they should be
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:34 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:33 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-omega in codfw: maintenance
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 19:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:29 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-omega in codfw: maintenance
* 19:27 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=k8s-ingress-aux-ro
* 19:26 ryankemper: [Cirrus] nevermind, that's just the cookbook assuming the DNS record should exist, which it doesn't because chi/psi/omega all share `search.svc.$DC.wmnet`
* 19:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 19:24 ryankemper: [Cirrus] `dns.resolver.NoAnswer: The DNS response does not contain an answer to the question: search-psi.svc.eqiad.wmnet` checking briefly if this is real failure or just some TTL wonkiness
* 19:23 ryankemper@cumin2003: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool search-psi in codfw: maintenance
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:20 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 19:18 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 19:18 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search-psi in codfw: maintenance
* 19:17 ryankemper: [Cirrus] codfw chi (big cluster) repooled; metrics are already improving, I see poolcounter rejections dropping significantly
* 19:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool search in codfw: maintenance
* 19:17 cdanis@cumin1004: conftool action : set/ttl=300; selector: name=codfw
* 19:13 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:12 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool search in codfw: maintenance
* 19:11 ryankemper: [Cirrus] Repooling codfw, chi first followed by the small clusters
* 19:11 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=(kartotherian{{!}}tegola-vector-tiles)
* 19:07 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 19:02 jhuneidi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] (duration: 10m 15s)
* 18:57 jhuneidi@deploy1003: daniel, jhuneidi: Continuing with deployment
* 18:56 jhuneidi@deploy1003: daniel, jhuneidi: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 18:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 18:52 jhuneidi@deploy1003: Started scap sync-world: Backport for [[gerrit:1344753{{!}}REST: restore PageContentHelper::checkAccess (fix live breakage)]]
* 18:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-scholarly in codfw: maintenance
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 18:48 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:48 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 18:44 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-scholarly in codfw: maintenance
* 18:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-internal-main in codfw: maintenance
* 18:42 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-swift,name=codfw
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:42 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-query,name=codfw
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 18:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 18:39 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:39 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-internal-main in codfw: maintenance
* 18:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wcqs in codfw: maintenance
* 18:37 herron@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=thanos-web.*,name=codfw
* 18:36 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:34 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wcqs in codfw: maintenance
* 18:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-scholarly in codfw: maintenance
* 18:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:30 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:28 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-scholarly in codfw: maintenance
* 18:25 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in codfw: maintenance
* 18:25 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 18:20 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in codfw: maintenance
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:19 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 18:18 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 18:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:15 ryankemper: [WDQS] Preparing to repool codfw WDQS shortly; it's been operating single DC so this second DC should restore proper service availability
* 18:13 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 18:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:11 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_ulsfo - 7.1.1-2~bpo13+wmf3 ()
* 18:10 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:04 taavi@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions (duration: 109m 23s)
* 18:04 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:02 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 18:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 18:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 17:59 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:59 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool codfw [reason: no reason specified, no task ID specified]
* 17:58 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 17:54 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:54 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 17:52 cdanis@cumin1004: conftool action : set/pooled=false; selector: name=codfw,dnsdisc=mwdebug.*
* 17:52 swfrench@cumin1004: conftool action : set/pooled=false; selector: dnsdisc=mwdebug.*,name=codfw
* 17:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-.*-ro
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apus
* 17:47 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mwdebug.*
* 17:47 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:44 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=swift
* 17:42 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=config-master{{!}}device-analytics{{!}}echostore{{!}}helm-charts{{!}}k8s-ingress-wikikube-ro{{!}}linkrecommendation{{!}}mathoid{{!}}restbase{{!}}restbase-async{{!}}rest-gateway-ro{{!}}mobileapps{{!}}mwdebug.*{{!}}push-notifications{{!}}recommendation-api{{!}}releases{{!}}wikifeeds
* 17:38 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-upload_magru and not P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:37 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on A:cp-text_magru and not P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 17:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 17:32 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=citoid{{!}}zotero
* 17:30 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=apertium{{!}}schema{{!}}termbox{{!}}proton{{!}}cxserver
* 17:22 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 17:19 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=thumbor
* 17:18 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=shellbox.*
* 17:17 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=urldownloader
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 17:05 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1001.*
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:55 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:54 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 16:49 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-next-ro
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 16:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 16:40 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:37 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-web-ro
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:33 cdanis@cumin1004: conftool action : set/pooled=true; selector: name=codfw,dnsdisc=mw-api-int-ro
* 16:33 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1001.eqiad.wmnet with OS trixie
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:23 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 16:20 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:19 hnowlan@cumin1004: END (FAIL) - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches (exit_code=99) for datacenter switchover from eqiad to codfw
* 16:15 taavi@deploy1003: Locking from deployment [ALL REPOSITORIES]: locked for re-pooling codfw for read traffic, contact SRE for equestions
* 16:14 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:14 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] (duration: 33m 52s)
* 16:08 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1001.eqiad.wmnet with reason: host reimage
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 16:03 swfrench@cumin1004: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo
* 16:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 16:01 hnowlan@cumin1004: START - Cookbook sre.switchdc.mediawiki.00-optional-warmup-caches for datacenter switchover from eqiad to codfw
* 16:01 swfrench@cumin1004: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo
* 16:01 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 16:00 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:57 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1017.eqiad.wmnet
* 15:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1001.eqiad.wmnet with OS trixie
* 15:51 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3005.*
* 15:51 swfrench-wmf: begin rolling restarts of confds in eqsin, codfw, ulsfo to reflect etcd SRV record changes
* 15:47 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1017.eqiad.wmnet
* 15:40 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344711{{!}}AbuseReview: Enable on enwiki (T439149)]], [[gerrit:1344693{{!}}Sync wmf/1.47.0-wmf.20 with wmf/1.47.0-wmf.21 for vandalism alpha (T438467)]]
* 15:35 vgutierrez@dns1004: END - running authdns-update
* 15:33 vgutierrez@dns1004: START - running authdns-update
* 15:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] (duration: 12m 33s)
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:30 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-ssd: apply
* 15:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:27 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 15:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 15:25 dreamyjazz@deploy1003: kharlan, dreamyjazz: Continuing with deployment
* 15:24 dreamyjazz@deploy1003: kharlan, dreamyjazz: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 15:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344694{{!}}EventMapper::fetchByPage: Allow filtering by type (T438031)]]
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3005.esams.wmnet with OS trixie
* 15:12 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox
* 15:06 vgutierrez@dns1004: END - running authdns-update
* 15:04 vgutierrez@dns1004: START - running authdns-update
* 15:03 vgutierrez@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns2.*,service=authdns-update
* 14:59 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] (duration: 32m 20s)
* 14:57 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7011.*
* 14:54 dkertesz@cumin1004: conftool action : set/pooled=yes; selector: name=cp7001.*
* 14:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:51 dkertesz: repooling cp7001{{!}}7011 after successful testing ([[phab:T343000|T343000]])
* 14:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:47 kharlan@deploy1003: kharlan: Continuing with deployment
* 14:46 kharlan@deploy1003: kharlan: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:43 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=eqiad,name=dse-k8s-worker1015.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:40 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:40 btullis@puppetserver1001: conftool action : set/pooled=yes; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1016.eqiad.wmnet
* 14:39 btullis@puppetserver1001: conftool action : set/weight=10; selector: service=kubesvc,cluster=dse-k8s,dc=codfw,name=dse-k8s-worker1015.eqiad.wmnet
* 14:39 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:39 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3005.esams.wmnet with reason: host reimage
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 14:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 14:34 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7011.*
* 14:33 dkertesz@cumin1004: conftool action : set/pooled=no; selector: name=cp7001.*
* 14:32 dkertesz: depooling cp7001{{!}}7011 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344222 (context: https://phabricator.wikimedia.org/T343000)
* 14:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 14:30 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1*.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1344684{{!}}AbuseReview: Add local CheckUsers to vandalism alpha test (T438467)]], [[gerrit:1344677{{!}}AbuseReview: Inidicate if the queue hides recent edits (T438235)]]
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:26 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:26 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:22 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:22 elukey: elukey@rdb2013:/srv/redis/appendonlydir$ sudo -u redis redis-check-aof --fix rdb2013-6380.aof.22039.incr.aof
* 14:21 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:21 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5025,5026].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:20 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 14:19 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:17 moritzm: installing Bird security updates
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 14:13 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:13 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:11 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:09 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3005.esams.wmnet with OS trixie
* 14:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir4004.*
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 14:07 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:07 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] (duration: 12m 27s)
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:05 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 14:03 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 urbanecm@deploy1003: migr, urbanecm: Continuing with deployment
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 14:01 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:01 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 14:00 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: migr, urbanecm: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 13:55 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:55 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:54 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344662{{!}}fix(AccountSetup): ensure TestKitchen knows about new user in CentralAuth redirect (T436872)]]
* 13:40 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] (duration: 10m 38s)
* 13:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 13:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:35 moritzm: installing nghttp2 security updates
* 13:35 awight@deploy1003: awight: Continuing with deployment
* 13:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir4004.ulsfo.wmnet with reason: host reimage
* 13:33 awight@deploy1003: awight: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1344246{{!}}Fixes failing edge when page is missing and entity usage remain. Updating ReallyDoQuery to function like an inner join. (T437687)]]
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'sync'.
* 13:26 elukey@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 13:25 elukey@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 13:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 13:18 mlitn@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] (duration: 12m 25s)
* 13:13 mlitn@deploy1003: mfossati, mlitn: Continuing with deployment
* 13:10 mlitn@deploy1003: mfossati, mlitn: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]] synced to the testservers (see https://wiki
* 13:08 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir4004.ulsfo.wmnet with OS trixie
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 mlitn@deploy1003: Started scap sync-world: Backport for [[gerrit:1344617{{!}}Instrument five-arm image carousel retest (T431362)]], [[gerrit:1344619{{!}}Wire image carousel retest instrumentation (T431362)]], [[gerrit:1344627{{!}}ThumbExtractor: trim nbsp and dangling colons from caption text (T435672)]], [[gerrit:1344630{{!}}ThumbExtractor: exclude lead infobox images from the carousel (T438907)]]
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:06 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 13:02 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:01 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=cp2059.*
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:52 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 12:34 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 12:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:21 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 11:13 hnowlan: restarted restbase on restbase2029
* 11:04 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[5031,5032].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:50 hnowlan: deleting stuck mw-web pods in eqiad
* 10:45 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] (duration: 10m 09s)
* 10:44 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:42 vgutierrez@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=1) rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:40 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:39 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:36 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1080.eqiad.wmnet with reason: provision
* 10:35 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344621{{!}}AbuseReview: Let specific users and suppressors see vandalism tag (T438860)]]
* 10:34 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:32 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] (duration: 10m 34s)
* 10:29 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 10:26 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:25 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:23 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1079.eqiad.wmnet with reason: provision
* 10:22 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 10:21 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344281{{!}}WikimediaAntiAbuse: Enable likely vandalism classifier on testwiki (T438860)]]
* 10:17 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall. (duration: 653m 55s)
* 10:17 hnowlan@deploy1003: Forcefully removing global lock: Unlocking scap after restoration of power in codfw
* 10:12 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 10:11 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:08 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_eqsin and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 10:03 moritzm: installing apr-util security updates
* 09:46 moritzm: installing bind9 security updates (client-side tools/libs only)
* 09:40 vgutierrez@puppetserver1001: conftool action : set/pooled=no; selector: name=cirrussearch1120.eqiad.wmnet
* 09:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:27 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 09:26 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B13
* 09:25 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B13
* 09:23 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:22 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:21 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:10 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker1016.eqiad.wmnet
* 09:05 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:04 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker1016.eqiad.wmnet
* 09:01 XioNoX: asw1-b13-drmrs> request system reboot - [[phab:T437984|T437984]]
* 09:00 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.reboot-runner (exit_code=0) rolling reboot on A:gitlab-runner
* 09:00 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B13
* 08:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:58 moritzm: installing node-lodash security updates
* 08:56 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B13
* 08:55 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1078.eqiad.wmnet with reason: provision
* 08:54 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1078.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 08:49 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'pool' for drmrs rack B12
* 08:47 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'pool' for drmrs rack B12
* 08:46 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 19 hosts with reason: Switches upgrade
* 08:46 moritzm: uploaded debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt
* 08:45 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b13-drmrs,asw1-b13-drmrs IPv6,asw1-b13-drmrs.mgmt with reason: Switch upgrade
* 08:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:37 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1015.eqiad.wmnet
* 08:31 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1015.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 08:22 XioNoX: asw1-b12-drmrs> request system reboot - [[phab:T437984|T437984]]
* 08:20 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for drmrs rack B12
* 08:13 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for drmrs rack B12
* 08:06 jelto@cumin1004: START - Cookbook sre.gitlab.reboot-runner rolling reboot on A:gitlab-runner
* 08:02 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-b12-drmrs,asw1-b12-drmrs IPv6,asw1-b12-drmrs.mgmt with reason: Switch upgrade
* 07:53 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 20 hosts with reason: Switches upgrade
* 07:52 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:52 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: switch upgrade, [[phab:T437984|T437984]]]
* 07:48 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:19 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: version upgrade
* 07:16 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:06 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: version upgrade
* 07:02 jelto@cumin1004: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:51 jelto@cumin1004: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: version upgrade
* 06:41 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production ([[phab:T437213|T437213]])
* 06:41 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/machinetranslation: apply
* 06:39 kart_: staging: Update machinetranslation/MinT to 2026-09-21-112314-production
* 06:38 kartik@deploy1003: helmfile [staging] START helmfile.d/services/machinetranslation: apply
* 06:07 ayounsi@cumin1004: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lsw1-e5-codfw
* 06:06 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 05:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:21 ryankemper@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster search_codfw: Restart codfw following today's power incident to ensure we return to our full expected state - ryankemper@cumin2003 - [[phab:T439010|T439010]]
* 01:19 ryankemper: [Cirrus] Reverted `node_concurrent_recoveries` to 5 from 10, now that we're back to green
* 01:16 ryankemper: [Cirrus] With the restart of `cirrussearch2115`, the codfw cluster has officially reached green status!!! Still working on full verification, but we're almost done here
* 01:14 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2115.codfw.wmnet with reason: Codfw survivor recovery on 2115; temporary chi red expected ([[phab:T439010|T439010]])
* 01:11 brett@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cp2059.codfw.wmnet with reason: failing services but not in service yet
* 01:10 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2109.codfw.wmnet with reason: Codfw survivor recovery on 2109; temporary chi red expected ([[phab:T439010|T439010]])
* 01:04 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2104.codfw.wmnet with reason: Codfw survivor recovery on 2104; temporary chi red expected ([[phab:T439010|T439010]])
* 01:03 ryankemper: [Cirrus] grr, I'd missed some hosts. restarting the last few dangling ones, we're really close to back to green, prob 3-ish more hosts
* 00:40 ryankemper: [Cirrus] Great news, we briefly dipped red (same as previous restarts) but went back to yellow almost immediately. AFAICT election went fine, still checking though
* 00:38 ryankemper: [Cirrus] Preparing to restart cirrussearch2084 (active cluster manager). With luck, this should restore updater availability (and general cluster green status, after some reshuffling)
* 00:35 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on 55 hosts with reason: Codfw chi elected-manager recovery on 2084; expected brief failover and red state ([[phab:T439010|T439010]])
* 00:10 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp7011.*
* 00:05 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 00:00 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
== 2026-09-23 ==
* 23:58 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 23:56 brett: Switching acme-chief primary from codfw to eqiad - [[phab:T439010|T439010]]
* 23:54 brett@cumin1004: END (FAIL) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=1) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:49 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7011.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 brett@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-varnish (exit_code=0) rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:48 ryankemper: [Cirrus] Every host except 2084, which is the current elected chi master, has now been restarted, and shard recoveries healed accordingly. AFAICT we will not be able to revive the updater until we restart this host. Pausing for a few mins to mull things over and get my bearings though, because this restart would be higher-touch than the previous ones
* 23:38 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2108.codfw.wmnet with reason: Codfw survivor recovery on 2108; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:38 brett@cumin1004: START - Cookbook sre.cdn.roll-upgrade-varnish rolling upgrade of Varnish on P<nowiki>{</nowiki>cp7001.magru.wmnet<nowiki>}</nowiki> and A:cp - 7.1.1-2~bpo13+wmf3 ()
* 23:35 brett: import varnish 7.1.1-2~bpo13+wmf3 into trixie-wikimedia ([[phab:T438293|T438293]])
* 23:34 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2107.codfw.wmnet with reason: Codfw survivor recovery on 2107; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:27 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2085.codfw.wmnet with reason: Codfw survivor recovery on 2085; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:23 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: No deployments please, as we're still cleaning up from the codfw power incident [[phab:T439010|T439010]]. Thursday UTC morning at the earliest, but please ask SRE oncall.
* 23:23 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 121m 40s)
* 23:20 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2072.codfw.wmnet with reason: Codfw survivor recovery on 2072; sequential chi and psi restarts ([[phab:T439010|T439010]])
* 23:09 ryankemper: [Cirrus] rolling cirrussearch2086 next
* 23:08 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2086.codfw.wmnet with reason: Codfw survivor recovery on 2086; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 23:01 ryankemper: [Cirrus] Doing cirrussearch2114 next
* 22:59 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2114.codfw.wmnet with reason: Codfw survivor recovery on 2114; sequential chi and omega restarts ([[phab:T439010|T439010]])
* 22:44 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2106.codfw.wmnet with reason: Codfw chi survivor recovery on 2106; temporary red expected ([[phab:T439010|T439010]])
* 22:29 ryankemper: [Cirrus] proceeding with manual restart of cirrussearch2105; red status expected, hopefully brief but we'll see
* 22:28 ryankemper@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch2105.codfw.wmnet with reason: Codfw chi recovery canary on 2105; temporary service interruption expected ([[phab:T439010|T439010]])
* 22:24 ryankemper: [Cirrus] s/expected/expect
* 22:23 ryankemper: [Cirrus] Alright, I'm getting increasingly convinced that there's no way to restore healthy cluster state without inevitably having to restart sole-shard-holder hosts, which will put the cluster into red status. going to start with just `cirrussearch2105`; I expected red status. silencing alerts first so I don't blow out the channel
* 22:08 ryankemper: [Cirrus] (to be clear the cluster is not serving live traffic, but if I can avoid red I will)
* 22:08 ryankemper: [Cirrus] updater still failing in codfw cirrussearch; i've restarted the directly-impacted hosts but not the others. some bulk updates appear to be getting rejected, going to do some targeted restarts and assess impact before considering a broader operation. first up is `cirrussearch2071.codfw.wmnet` which is not the sole holder of any shards therefore should not plunge the cluster into red status
* 21:49 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 21:22 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 21:22 rzl@deploy1003: Unlocked for deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]] (duration: 51m 29s)
* 21:21 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 21:18 Emperor: ceph mgr fail on apus-be2005
* 21:18 Emperor: reset-failed then restart ceph-mon on moss-be2003
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2235].codfw.wmnet with reason: needs fixing
* 21:08 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2234].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2233].codfw.wmnet with reason: needs fixing
* 21:07 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2160,2232].codfw.wmnet with reason: needs fixing
* 20:57 ryankemper: [Cirrus] cirrussearch codfw back to yellow status. active shard pct = 94.51%
* 20:55 ryankemper: [Cirrus] Bump codfw cirrussearch shard recoveries from 5 to 10; cluster not serving live traffic so I'm hoping we have headroom to recover faster
* 20:49 swfrench@dns1004: END - running authdns-update
* 20:46 swfrench@dns1004: START - running authdns-update
* 20:41 ryankemper: [Cirrus] Been restarting all impacted codfw opensearch hosts one at a time (they didn't rejoin the cluster naturally)
* 20:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 20:38 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ncredir5004.eqsin.wmnet with OS trixie
* 20:30 rzl@deploy1003: Locking from deployment [ALL REPOSITORIES]: incident recovery in progress [[phab:T439010|T439010]]
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:27 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:06 dzahn@dns1004: END - running authdns-update
* 20:03 dzahn@dns1004: START - running authdns-update
* 19:52 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5004.eqsin.wmnet with OS trixie
* 19:34 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp2059.codfw.wmnet with OS trixie
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:33 volans: rebooting arclamp2001.codfw.wmnet
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:20 sukhe@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 979 hosts with reason: power is still coming back on
* 19:17 taavi@dns1004: END - running authdns-update
* 19:14 taavi@dns1004: START - running authdns-update
* 19:10 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 19:10 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 19:10 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6001.*
* 19:08 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:59 sukhe@cumin1004: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 6:00:00 on 980 hosts with reason: power is still coming back on
* 18:58 taavi@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) gerrit.discovery.wmnet on all recursors
* 18:58 taavi@cumin1004: START - Cookbook sre.dns.wipe-cache gerrit.discovery.wmnet on all recursors
* 18:50 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit2003.wikimedia.org
* 18:45 sukhe@dns1004: END - running authdns-update
* 18:43 sukhe@dns1004: START - running authdns-update
* 18:43 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: dc=codfw,cluster=dnsbox,service=authdns-update
* 18:42 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:42 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: END (FAIL) - Cookbook sre.gerrit.localbackup (exit_code=99) Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 dzahn@cumin2003: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit2003.wikimedia.org
* 18:40 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.localbackup (exit_code=0) Prepare local backup on: gerrit1003.wikimedia.org
* 18:38 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:38 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet _etcd-client-ssl._tcp.ulsfo.wmnet _etcd-client-ssl._tcp.codfw.wmnet on all recursors
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit1003.wikimedia.org
* 18:36 taavi@cumin1004: END (PASS) - Cookbook sre.gerrit.read-only-toggle (exit_code=0) from gerrit2003.wikimedia.org
* 18:36 taavi@cumin1004: START - Cookbook sre.gerrit.read-only-toggle from gerrit2003.wikimedia.org
* 18:30 taavi@cumin1004: START - Cookbook sre.gerrit.localbackup Prepare local backup on: gerrit1003.wikimedia.org
* 18:29 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:28 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:14 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 18:08 sukhe@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) idp.wikimedia.org on all recursors
* 18:08 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache idp.wikimedia.org on all recursors
* 18:05 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:05 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:03 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.eqsin.wmnet on all recursors
* 18:02 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:02 cdanis@cumin1004: START - Cookbook sre.dns.wipe-cache _etcd-client-ssl._tcp.ulsfo.wmnet on all recursors
* 18:01 cdanis@dns1005: END - running authdns-update
* 17:58 cdanis@dns1005: START - running authdns-update
* 17:57 vriley@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:54 taavi@dns1004: END - running authdns-update
* 17:53 vriley@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:51 taavi@dns1004: START - running authdns-update
* 17:46 taavi@dns1004: END - running authdns-update
* 17:43 taavi@dns1004: START - running authdns-update
* 17:37 vriley@cumin1004: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:35 rzl@cumin1004: START - Cookbook sre.discovery.datacenter pool all active/active services in eqiad: maintenance - [[phab:T439010|T439010]]
* 17:35 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6001.drmrs.wmnet with OS trixie
* 17:35 cdanis@cumin1004: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:35 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:24 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:23 sukhe@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool codfw [reason: no reason specified, no task ID specified]
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:21 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:18 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:17 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:16 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:14 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049.codfw.wmnet
* 17:11 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp2049
* 17:10 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cp2059.codfw.wmnet with reason: host reimage
* 17:07 mutante: cloudcontrol2005-dev, cloudcontrol2006-dev, cloudcontrol2010-dev: restart zookeeper, enabled logging (/var/log/zookeeper/zookeeper.log) after gerrit:1342354
* 17:02 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:59 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6001.drmrs.wmnet with reason: host reimage
* 16:51 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:51 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp2059.codfw.wmnet with OS trixie
* 16:48 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:39 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 16:35 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 16:34 dzahn@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1005.eqiad.wmnet with OS trixie
* 16:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:29 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6001.drmrs.wmnet with OS trixie
* 16:10 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 16:10 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:55 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:54 moritzm: installing cjose security updates
* 15:54 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir7004.*
* 15:53 dancy@deploy1003: Finished scap sync-world: testing (duration: 07m 06s)
* 15:52 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:52 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp2059.codfw.wmnet with OS trixie
* 15:51 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 15:46 dancy@deploy1003: Started scap sync-world: testing
* 15:43 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host cp2059.codfw.wmnet with OS trixie
* 15:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir7004.magru.wmnet with OS trixie
* 15:42 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 15:41 sukhe: homer "lsw1-e4-codfw.*" commit 'pending from cookbook'
* 15:41 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.c7/c/c7/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(248_of_580).jpg' codfw:wikipedia-commons-local-public.c7/c/c7 [[phab:T438961|T438961]]
* 15:39 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.rename (exit_code=0) from sretest2013 to cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cp2059
* 15:38 sukhe@cumin1004: START - Cookbook sre.network.configure-switch-interfaces for host cp2059
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cp2059 on all recursors
* 15:38 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.a9/a/a9/Ğāmi‛_al-tavārīḫ._Rašīd_al-Dīn_Fazl-ullāh_Hamadānī_-_btv1b8427170s_(182_of_597).jpg' codfw:wikipedia-commons-local-public.a9/a/a9/ [[phab:T438961|T438961]]
* 15:38 sukhe@cumin1004: START - Cookbook sre.dns.wipe-cache cp2059 on all recursors
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:38 sukhe@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:37 sukhe@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Renaming sretest2013 to cp2059 - sukhe@cumin1004"
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:36 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:35 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.4d/4/4d/Kamāl_al-Dīn_Ḥusayn_b._ʿAlī_Bayhaqī_Sabzavārī_Vā‛iẓ_Kāšifī_._Anvār-i_Suhaylī_-_btv1b10515885n_(142_of_580).jpg' codfw:wikipedia-commons-local-public.4d/4/4d [[phab:T438961|T438961]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 15:35 mutante: zuul1005 - reimage - should not have had nftables on it before [[phab:T438786|T438786]]
* 15:35 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 15:34 dzahn@cumin2003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 15:34 sukhe@cumin1004: START - Cookbook sre.dns.netbox
* 15:33 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:32 Emperor: rclone copy --no-update-modtime --checksum --config /etc/swift/rclone.conf 'eqiad:wikipedia-commons-local-public.41/4/41/ĞAVĀMI‛_al-ḤIKĀYĀT_VA_LAVĀMI‛_al-RIVĀYĀT._Sadīd_al-Dīn_Muḥ._b._Muḥ._b._Yaḥyà_‛Awfī_Buhārī_Ḥanafī._-_btv1b525129105_(033_of_524).jpg' codfw:wikipedia-commons-local-public.41/4/41 [[phab:T438961|T438961]]
* 15:23 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 15:23 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_ulsfo and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:21 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2013.codfw.wmnet with OS trixie
* 15:21 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 15:21 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 15:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 15:19 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 15:18 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:14 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir7004.magru.wmnet with reason: host reimage
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-ro,name=eqiad
* 15:12 jayme@deploy1003: conftool action : set/pooled=true; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:12 moritzm: removed buster-wikimedia and all related components from apt.wikimedia.org following the merge of https://gerrit.wikimedia.org/r/c/operations/puppet/+/1247618
* 15:06 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 15:02 dancy@deploy1003: Installation of scap version "4.292.0" completed for 3 hosts
* 15:02 sukhe@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 15:02 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 15:01 jayme@deploy1003: conftool action : set/pooled=false; selector: dnsdisc=mw-web-next-ro,name=eqiad
* 15:01 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 15:01 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:00 dancy@deploy1003: Installing scap version "4.292.0" for 3 host(s)
* 15:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 14:58 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 14:57 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 14:56 sukhe@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2013.codfw.wmnet with reason: host reimage
* 14:55 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:55 jayme@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) depool mw-web-ro in eqiad: maintenance
* 14:55 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:54 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route depool mw-web-ro in eqiad: maintenance
* 14:54 jayme@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 14:54 jayme@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 14:53 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:53 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:53 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:47 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:47 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:47 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:46 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 14:45 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 14:45 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:45 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 14:43 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:43 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:42 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:42 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 12s)
* 14:42 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 14:42 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir7004.magru.wmnet with OS trixie
* 14:40 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 14:40 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from codfw to eqiad for section test-s4
* 14:39 moritzm: upload debuerreotype 0.15-1.1+wmf13u1 to component/main from trixie-wikimedia [[phab:T438866|T438866]]
* 14:38 gengh@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 gengh@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]] (duration: 14m 34s)
* 14:37 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-upload_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:37 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:36 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:36 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 14:28 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 14:28 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir1002.*
* 14:26 gengh@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:26 gengh@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:24 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:23 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:23 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1344292{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]], [[gerrit:1344293{{!}}feat(AddLink): Do not resuggest an already reviewed page (T429417)]]
* 14:17 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir1002.eqiad.wmnet with OS trixie
* 14:10 gengh@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:09 gengh@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 14:07 ebernhardson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search: apply
* 13:58 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:56 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 13:53 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir1002.eqiad.wmnet with reason: host reimage
* 13:38 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on A:cp-text_magru and not P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:37 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:37 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:35 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir1002.eqiad.wmnet with OS trixie
* 13:30 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 13:29 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs-test1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:25 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:24 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:24 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 13:18 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 13:13 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from codfw to eqiad for section test-s4
* 13:10 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from codfw to eqiad for section test-s4
* 13:09 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.finalize (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 13:04 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.finalize for the switch from eqiad to codfw for section test-s4
* 12:57 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:57 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:56 awight: manually run puppet agent
* 12:56 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:56 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:55 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:55 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:45 awight: add seanleong-wmde to deployment-prep
* 12:44 cmooney@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-eqiad,ssw1-d[1,8]-eqiad with reason: re-rack ssw1-a1-eqiad
* 12:39 cwilliams@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section test-s4
* 12:39 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:38 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:34 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] (duration: 33m 33s)
* 12:33 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:33 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:32 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 12:32 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:30 brouberol@cumin1004: END (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 12:30 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 12:29 cwilliams@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section test-s4
* 12:22 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:21 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:15 cdanis@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:15 cdanis@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool eqiad [reason: no reason specified, no task ID specified]
* 12:01 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343982{{!}}AbuseReview: Add warning indicating alpha test to vandalism queue (T438467)]]
* 11:51 Dreamy_Jazz: Deployed patch for [[phab:T438729|T438729]]
* 11:31 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:28 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:27 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:25 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 10:38 jayme: sudo confctl --quiet --object-type discovery select 'dnsdisc=mw-web-ro' set/ttl=10 - [[phab:T438896|T438896]]
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:25 blake@deploy1003: Finished scap sync-world: Upsize mw-web [[phab:T438896|T438896]] (duration: 04m 20s)
* 10:22 blake@deploy1003: Started scap sync-world: Upsize mw-web [[phab:T438896|T438896]]
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 10:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 10:04 ayounsi@cumin1004: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:59 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:59 ayounsi@cumin1004: START - Cookbook sre.deploy.python-code netbox to netbox-dev2003.codfw.wmnet with reason: Add netbox-bgp and update wheelson netbox-next - ayounsi@cumin1004
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 09:58 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:58 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:57 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:55 brouberol@cumin1004: DONE (PASS) - Cookbook sre.ceph.remove-osd (exit_code=0)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:54 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:53 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:52 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 09:51 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:51 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:50 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:44 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 09:43 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-wdqs100[1-3].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 09:38 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 09:34 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:45 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:44 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:41 brouberol@cumin1004: DONE (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: END (FAIL) - Cookbook sre.ceph.remove-osd (exit_code=99)
* 08:27 brouberol@cumin1004: START - Cookbook sre.ceph.remove-osd
* 08:25 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:24 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:13 tappof@deploy1003: Finished scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8] (duration: 12m 52s)
* 08:05 moritzm: installing grub2 bugfix updates on Bookworm hosts
* 08:04 tappof@deploy1003: Started scap sync-world: [[phab:T432444|T432444]] - Provision kafka-logging100[6-8]
* 08:00 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 07:59 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 07:59 moritzm: installing giflib security updates
* 07:58 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 07:58 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 07:29 moritzm: installing python-idna security updates
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 39s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool wdqs-main in eqiad: maintenance
* 00:46 ryankemper: [WDQS] [[phab:T435443|T435443]] Restore eqiad wdqs-main; wdqs was unable to keep up with traffic with only one datacenter. sadly this will continue to be the case until wdqsv2 is ready to switch backend architecture
* 00:45 ryankemper@cumin2003: START - Cookbook sre.discovery.service-route pool wdqs-main in eqiad: maintenance
== 2026-09-22 ==
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 23:23 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:23 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 23:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 22:45 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:45 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 22:36 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:30 ryankemper: [WDQS] codfw wdqs-main is struggling under the switchover load, fiddling with some auto-restart knobs to see if it helps or hurts
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 22:06 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 22:06 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 21:58 rzl@deploy1003: Finished scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] (duration: 13m 43s)
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 21:57 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 21:53 rzl@deploy1003: rzl: Continuing with deployment
* 21:51 rzl@deploy1003: rzl: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:47 rzl@deploy1003: Started scap sync-world: https://gerrit.wikimedia.org/r/1339694 [[phab:T437403|T437403]]
* 21:25 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 01m 09s)
* 21:24 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 21:24 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 24m 20s)
* 21:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 21:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 21:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 21:05 sukhe@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2013.codfw.wmnet with OS trixie
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:59 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566] (duration: 00m 30s)
* 20:59 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (thin): Regular analytics weekly train THIN [analytics/refinery@58c93566]
* 20:55 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566] (duration: 06m 59s)
* 20:48 aqu@deploy1003: Started deploy [analytics/refinery@58c9356]: Regular analytics weekly train [analytics/refinery@58c93566]
* 20:46 aqu@deploy1003: Finished deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566] (duration: 00m 40s)
* 20:45 bking@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 20:45 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] (duration: 09m 53s)
* 20:45 aqu@deploy1003: Started deploy [analytics/refinery@58c9356] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@58c93566]
* 20:44 bking@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 20:44 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:43 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:40 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 20:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 20:40 sbisson@deploy1003: sbisson: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:35 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342285{{!}}Keep Article Guidance on where it is on today (T433293)]]
* 20:33 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] (duration: 13m 35s)
* 20:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:28 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Continuing with deployment
* 20:24 ebernhardson@deploy1003: ebernhardson, codenamenoreste: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1342825{{!}}eswiki: Add abusefilter-access-protected-vars to abusefilter user group (T436652)]]
* 20:17 ebernhardson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] (duration: 10m 29s)
* 20:15 bking@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:13 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir2002.*
* 20:12 bking@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:12 ebernhardson@deploy1003: ebernhardson: Continuing with deployment
* 20:11 ebernhardson@deploy1003: ebernhardson: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:06 ebernhardson@deploy1003: Started scap sync-world: Backport for [[gerrit:1344014{{!}}cirrus: Send more_like traffic to eqiad]]
* 20:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 20:02 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 20:02 sukhe@cumin1004: START - Cookbook sre.hosts.reimage for host sretest2013.codfw.wmnet with OS trixie
* 19:59 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir2002.codfw.wmnet with OS trixie
* 19:44 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 19:42 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:42 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 19:38 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir2002.codfw.wmnet with reason: host reimage
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1016.eqiad.wmnet with OS trixie
* 19:34 jclark@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:33 jclark@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jclark@cumin1004"
* 19:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:24 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:23 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:22 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 19:21 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 19:19 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir2002.codfw.wmnet with OS trixie
* 19:19 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 19:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 19:15 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1016.eqiad.wmnet with reason: host reimage
* 19:01 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-codfw to update to opensearch 3.8.0
* 18:58 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 18:56 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 18:56 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:56 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 18:55 dancy@deploy1003: Installation of scap version "4.291.0" completed for 2 hosts
* 18:53 dancy@deploy1003: Installing scap version "4.291.0" for 2 host(s)
* 18:53 dancy@deploy1003: Installation of scap version "4.291.0" completed for 3 hosts
* 18:51 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 18:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:49 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 18:47 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:44 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:43 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:41 dancy@deploy1003: install-world aborted: (no justification provided) (duration: 00m 48s)
* 18:41 dancy@deploy1003: Installing scap version "4.291.0" for 3 host(s)
* 18:40 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:36 jhuneidi@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 18:35 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:35 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 18:18 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 18:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 18:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 18:07 ebernhardson: Rolling restart opensearch-semantic-search in dse-k8s-eqiad to update to opensearch 3.8.0
* 17:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] (duration: 10m 09s)
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1030
* 17:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1030
* 17:53 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1030
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1030.eqiad.wmnet 8.32.64.10.in-addr.arpa 8.0.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:53 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1030 - bking@cumin2003"
* 17:51 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Optimizer issues fixed
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Continuing with deployment
* 17:50 dreamyjazz@deploy1003: dreamyjazz, isaranto: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:47 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:46 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344040{{!}}fix(WikimediaAntiAbuse): use correct endpoint for LiftWing in eqiad]]
* 17:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:43 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1030
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 17:39 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:39 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1029
* 17:37 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1029
* 17:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1029
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1029.eqiad.wmnet 8.48.64.10.in-addr.arpa 8.0.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1029 - bking@cumin2003"
* 17:33 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:32 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1029
* 17:26 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:25 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:25 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 17:24 dzahn@dns1004: END - running authdns-update
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 17:24 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:24 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 17:22 dzahn@dns1004: START - running authdns-update
* 17:17 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:17 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 17:16 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:16 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 17:15 oblivian@puppetserver1001: conftool action : set/pooled=false; selector: dnsdisc=rest-gateway,name=codfw
* 17:10 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 17:09 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:09 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 17:06 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool db2218: Optimizer issues fixed
* 17:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 17:02 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059.codfw.wmnet
* 17:01 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 17:00 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:59 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] (duration: 13m 13s)
* 16:54 oblivian@cumin1004: END (FAIL) - Cookbook sre.discovery.service-route (exit_code=99) pool 2 services in eqiad: maintenance
* 16:51 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 16:50 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:48 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 2 services in eqiad: maintenance
* 16:46 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:45 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1344020{{!}}Enable AbuseReview on jawiki for likely PII (T438867)]]
* 16:42 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on db2218.codfw.wmnet with reason: fixing
* 16:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:41 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 16:33 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:33 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 16:26 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:26 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 16:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
* 16:19 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool 4 services in eqiad: maintenance
* 16:13 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool 4 services in eqiad: maintenance
* 16:04 sukhe@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 16:04 sukhe@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 15:58 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2218: optimizer issues
* 15:57 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool db2218: optimizer issues
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 15:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:55 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 15:50 moritzm: installing libhtml-parser-perl security updates
* 15:49 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:40 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool mw-web-ro in eqiad: maintenance
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:36 ayounsi@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:36 ayounsi@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cirrussearch1120 move vlan - ayounsi@cumin1004"
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route pool mw-web-ro in eqiad: maintenance
* 15:35 oblivian@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check mw-web-ro: maintenance
* 15:35 oblivian@cumin1004: START - Cookbook sre.discovery.service-route check mw-web-ro: maintenance
* 15:27 ayounsi@cumin1004: START - Cookbook sre.dns.netbox
* 15:22 slyngshede@cumin1004: END (PASS) - Cookbook sre.discovery.datacenter (exit_code=0) depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 15:19 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 15:18 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:18 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1120
* 15:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1120
* 15:14 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.move-vlan (exit_code=99) for host cirrussearch1120
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 15:11 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:11 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 15:04 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 15:03 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:03 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 15:01 dancy@deploy1003: Installation of scap version "4.290.0" completed for 3 hosts
* 14:59 dancy@deploy1003: Installing scap version "4.290.0" for 3 host(s)
* 14:55 slyngshede@cumin1004: START - Cookbook sre.discovery.datacenter depool all services in eqiad: Datacenter services switchover - [[phab:T435443|T435443]]
* 14:55 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:54 dancy@deploy1003: Installing scap version "4.290.0" for 155 host(s)
* 14:54 slyngshede@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:54 slyngshede@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool eqiad [reason: no reason specified, no task ID specified]
* 14:53 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:51 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cirrussearch1120.eqiad.wmnet with reason: migrate VLAN [[phab:T436571|T436571]]
* 14:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1120
* 14:47 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1120
* 14:42 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:42 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:36 cdobbins@cumin1004: END (FAIL) - Cookbook sre.hosts.rename (exit_code=93) from sretest2013 to cp2059
* 14:35 cdobbins@cumin1004: START - Cookbook sre.hosts.rename from sretest2013 to cp2059
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 14:25 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:25 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 14:17 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 14:15 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 14:15 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 13:59 btullis@cumin1004: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 13:57 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker10[02-28].eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 13:57 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:56 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:55 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 13:53 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.powercycle (exit_code=0) for host pki1002
* 13:51 elukey@cumin1004: START - Cookbook sre.hosts.powercycle for host pki1002
* 13:23 tappof@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 13:22 tappof@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 13:21 tappof@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 12:53 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:48 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-ctrl1001.eqiad.wmnet
* 12:44 marostegui: Stop mariadb on db2250:s5 [[phab:T437411|T437411]] [[phab:T437279|T437279]]
* 12:43 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2250.codfw.wmnet with reason: preparations
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 12:31 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:31 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 12:22 btullis@cumin1004: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 12:19 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] (duration: 33m 01s)
* 12:17 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:16 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:15 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:08 kharlan@deploy1003: kharlan: Continuing with deployment
* 12:06 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:54 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host (duration: 00m 54s)
* 11:54 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to production host
* 11:52 jnuche@deploy1003: Finished deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host (duration: 01m 01s)
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 11:52 btullis@cumin1004: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>dse-k8s-worker1001.eqiad.wmnet<nowiki>}</nowiki> and (A:dse-k8s-master-eqiad or A:dse-k8s-worker-eqiad)
* 11:52 jnuche@deploy1003: Started deploy [releng/jenkins-deploy@ddb3f1a] (releasing): [[phab:T435791|T435791]] to backup host
* 11:46 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343953{{!}}AbuseReview: Hide recently saved revisions from the vandalism queue (T438235)]]
* 11:41 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] (duration: 13m 46s)
* 11:34 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:33 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:27 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343960{{!}}AbuseReview: Hide Echo banner when user cannot see personal info (T438477)]]
* 11:24 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] (duration: 33m 09s)
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 11:24 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 11:19 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 11:13 kharlan@deploy1003: kharlan: Continuing with deployment
* 11:12 kharlan@deploy1003: kharlan: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:54 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:53 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:52 topranks: enable rule cache-upload/eqsin_originals_scraper_20260922
* 10:51 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1343952{{!}}AbuseReview: Allow interaction with verdict buttons on closed rows (T438808)]]
* 10:20 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2005.codfw.wmnet with OS trixie
* 10:13 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s1
* 10:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s1
* 10:10 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s4
* 10:10 gmodena@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 10:09 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s4
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 10:09 gmodena@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 10:08 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s8
* 10:06 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s8
* 10:06 moritzm: installing libcap2 security updates
* 10:05 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s7
* 10:03 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s7
* 10:02 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 10:02 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s3
* 10:01 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s3
* 10:00 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s2
* 09:58 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s2
* 09:58 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2005.codfw.wmnet with reason: host reimage
* 09:57 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s5
* 09:56 vgutierrez@cumin1004: END (PASS) - Cookbook sre.cdn.roll-upgrade-haproxy (exit_code=0) rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:55 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s5
* 09:53 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section s6
* 09:51 elukey: install spicerack 13.3.0 on cumin1004 and cumin2003
* 09:50 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section s6
* 09:47 elukey: uploaded spicerack_13.3.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:47 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: issues
* 09:46 marostegui@cumin1004: START - Cookbook sre.mysql.pool pool es1035: issues
* 09:44 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:44 vgutierrez@cumin1004: START - Cookbook sre.cdn.roll-upgrade-haproxy rolling upgrade of HAProxy on P<nowiki>{</nowiki>cp[7010,7016].*<nowiki>}</nowiki> and A:cp - 3.2.23 upgrade ([[phab:T438828|T438828]])
* 09:44 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:43 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:42 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:41 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:41 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2005.codfw.wmnet with OS trixie
* 09:40 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es7
* 09:39 vgutierrez: fetch haproxy 3.2.23 on thirdparty/haproxy32 for trixie (apt.wm.o) - [[phab:T438828|T438828]]
* 09:32 marostegui@cumin1004: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: issues
* 09:32 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:32 marostegui@cumin1004: START - Cookbook sre.mysql.depool depool es1035: issues
* 09:31 marostegui@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 8 hosts with reason: dc preparations
* 09:30 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:28 jelto@cumin1004: conftool action : set/pooled=inactive; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:28 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 09:26 btullis@dns1004: END - running authdns-update
* 09:24 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 09:23 btullis@dns1004: START - running authdns-update
* 09:23 jelto@cumin1004: conftool action : set/pooled=no; selector: name=wikikube-worker1152.eqiad.wmnet
* 09:20 jelto@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker1152.eqiad.wmnet with reason: hardware/networking issues
* 09:16 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es7
* 09:15 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section es6
* 09:14 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section es6
* 09:12 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x4
* 09:11 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x3
* 09:10 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'email' for AS: 52320
* 09:09 ayounsi@cumin1004: START - Cookbook sre.network.peering with action 'email' for AS: 52320
* 09:05 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x3
* 09:04 fceratto@cumin1004: END (PASS) - Cookbook sre.switchdc.databases.prepare (exit_code=0) for the switch from eqiad to codfw for section x1
* 09:02 fceratto@cumin1004: START - Cookbook sre.switchdc.databases.prepare for the switch from eqiad to codfw for section x1
* 08:58 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:55 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:52 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:49 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:45 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:45 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt
* 08:44 ayounsi@cumin1004: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 13 hosts
* 08:44 ayounsi@cumin1004: START - Cookbook sre.hosts.remove-downtime for 13 hosts
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 08:39 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 08:37 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 08:32 moritzm: installig zip security updates
* 08:30 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:29 XioNoX: asw1-bw27-esams> request system reboot - [[phab:T437984|T437984]]
* 08:28 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:27 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for esams rack BW27
* 08:26 jelto@cumin1004: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for alias: wikikube-worker-eqiad@eqiad
* 08:26 ayounsi@cumin1004: START - Cookbook sre.network.depool-rack with action 'depool' for esams rack BW27
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:22 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 08:18 moritzm: installing gst-plugins-base1.0 security updates
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 08:10 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:09 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:05 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:05 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:04 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 13 hosts with reason: Switch reboot
* 08:01 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:01 ayounsi@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on asw1-bw27-esams,asw1-bw27-esams IPv6,asw1-bw27-esams.mgmt with reason: Switch reboot
* 07:59 ayounsi@cumin1004: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:59 ayounsi@cumin1004: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: switch reboot, [[phab:T437984|T437984]]]
* 07:23 awight: UTC morning deployment window complete
* 07:22 awight@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] (duration: 17m 46s)
* 07:15 awight@deploy1003: seanleong-wmde, esanders, awight: Continuing with deployment
* 07:09 awight@deploy1003: seanleong-wmde, esanders, awight: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:05 awight@deploy1003: Started scap sync-world: Backport for [[gerrit:1343347{{!}}Config change for launch of stopping sending LL notifications. (T438463)]], [[gerrit:1313951{{!}}Change feedback URLs for EditCheck TextMatch on ruwiki (T426271)]]
* 07:02 moritzm: installing pyasn1 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.18 (duration: 02m 28s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]] (duration: 35m 52s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.21 refs [[phab:T438217|T438217]]
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 30s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-09-21 ==
* 22:11 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:10 rzl@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:08 rzl@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 22:07 rzl@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 22:06 rzl@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:05 rzl@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:18 maryum: Deployed security fix for [[phab:T437708|T437708]]
* 20:35 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] (duration: 15m 56s)
* 20:30 cjming@deploy1003: ameisenigel, cjming: Continuing with deployment
* 20:26 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:25 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:23 cjming@deploy1003: ameisenigel, cjming: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:19 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1343100{{!}}Disable wgMFCustomSiteModules on German Wikipedia (T403380)]]
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:02 sfaci@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:59 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:35 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) pool sessionstore in eqiad: sessionstore1005 repaired
* 18:32 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir5003.*
* 18:30 Emperor: repool eqiad sessionstore [[phab:T437915|T437915]]
* 18:30 mvernon@cumin1004: START - Cookbook sre.discovery.service-route pool sessionstore in eqiad: sessionstore1005 repaired
* 18:27 mvernon@cumin1004: END (PASS) - Cookbook sre.discovery.service-route (exit_code=0) check sessionstore: maintenance
* 18:27 mvernon@cumin1004: START - Cookbook sre.discovery.service-route check sessionstore: maintenance
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:25 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 18:24 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir5003.eqsin.wmnet with OS trixie
* 17:54 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:50 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir5003.eqsin.wmnet with reason: host reimage
* 17:40 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sessionstore1005.eqiad.wmnet with OS bookworm
* 17:30 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:29 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:26 jclark@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on sessionstore1005.eqiad.wmnet with reason: host reimage
* 17:12 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 17:00 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host sessionstore1005.eqiad.wmnet with OS bookworm
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:56 ebernhardson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 16:54 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir5003.eqsin.wmnet with OS trixie
* 16:46 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir6002.*
* 16:44 jclark@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:44 tappof@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on kafka-logging1003.eqiad.wmnet with reason: migrating to kafka-logging1006
* 16:36 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir6002.drmrs.wmnet with OS trixie
* 16:32 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:27 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:23 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:22 jclark@cumin1004: START - Cookbook sre.hosts.provision for host sessionstore1005.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 16:16 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:15 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:15 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 16:13 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:10 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 16:09 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir6002.drmrs.wmnet with reason: host reimage
* 16:01 cklimas@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [codfw] START helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifeeds: apply
* 16:00 cklimas@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifeeds: apply
* 16:00 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host registry2004.codfw.wmnet with OS trixie
* 15:55 cklimas@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifeeds: apply
* 15:54 cklimas@deploy1003: helmfile [staging] START helmfile.d/services/wikifeeds: apply
* 15:49 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:45 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] (duration: 12m 40s)
* 15:42 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:39 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 15:39 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir6002.drmrs.wmnet with OS trixie
* 15:38 elukey@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on registry2004.codfw.wmnet with reason: host reimage
* 15:36 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:33 slyngshede@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-api-ext: apply
* 15:32 slyngshede@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-api-ext: apply
* 15:32 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1343579{{!}}Fixes: '.action_context' should be string (T437122)]]
* 15:19 elukey@puppetserver1001: conftool action : set/pooled=false; selector: name=registry2004.*
* 15:18 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host registry2004.codfw.wmnet with OS trixie
* 15:16 cdobbins@cumin1004: conftool action : set/pooled=yes; selector: name=ncredir3006.*
* 15:11 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 15:07 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-web: apply
* 15:07 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-web: apply
* 15:03 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ncredir3006.esams.wmnet with OS trixie
* 15:01 slyngshede@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-api-ext: apply
* 15:01 slyngshede@deploy1003: helmfile [codfw] START helmfile.d/services/mw-api-ext: apply
* 14:47 elukey@cumin1004: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:39 cdobbins@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:36 elukey@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 14:34 cdobbins@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on ncredir3006.esams.wmnet with reason: host reimage
* 14:26 elukey@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:20 elukey@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:13 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] (duration: 08m 09s)
* 14:08 zabe@deploy1003: zabe: Continuing with deployment
* 14:08 zabe@deploy1003: zabe: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:07 cdobbins@cumin1004: START - Cookbook sre.hosts.reimage for host ncredir3006.esams.wmnet with OS trixie
* 14:05 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 14:01 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1343542{{!}}Move wbc_entity_usage to x1 for mediawikiwiki (T438716)]], [[gerrit:1343556{{!}}Set db explicitly to false for virtual-wikibase-entityusage]]
* 13:55 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] (duration: 12m 30s)
* 13:51 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Continuing with deployment
* 13:47 dreamyjazz@deploy1003: dreamyjazz, novemlinguae: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1004: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:45 cmooney@cumin1004: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add entries for new eqiad links - cmooney@cumin1004"
* 13:43 zabe: reconcile wbc_entity_usage from local cluster to x1 for mediawikiwiki # [[phab:T438716|T438716]]
* 13:43 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1337580{{!}}nlwiki: enable SecurePoll local elections (T434045)]]
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 13:41 cmooney@cumin1004: START - Cookbook sre.dns.netbox
* 13:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] (duration: 11m 40s)
* 13:36 samtar@deploy1003: samtar, tryvix1509: Continuing with deployment
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 13:33 samtar@deploy1003: samtar, tryvix1509: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:29 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1343319{{!}}arywiki: Create patroller and autopatrolled user groups (T438421)]]
* 13:22 mfossati@deploy1003: Finished scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] (duration: 14m 19s)
* 13:15 mfossati@deploy1003: mfossati: Continuing with deployment
* 13:14 mfossati@deploy1003: mfossati: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:10 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1063.eqiad.wmnet
* 13:07 mfossati@deploy1003: Started scap sync-world: Backport for [[gerrit:1343122{{!}}Let AA measure eligible readers w/o beta opt-in (T437076)]]
* 13:01 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 12:59 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1063.eqiad.wmnet
* 12:57 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:54 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ml-serve1016.eqiad.wmnet with OS trixie
* 12:54 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 12:53 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:51 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:48 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:36 XioNoX: delete BGP sessions to 15305 in Equinix Ashburn (peer leaving the IX)
* 12:30 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 12:28 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:27 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 12:23 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 12:05 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2005.codfw.wmnet
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:45 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:45 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2005.codfw.wmnet
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 11:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 11:42 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 11:41 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 11:40 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:40 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2004.codfw.wmnet
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 11:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-experiment-platform: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 11:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 11:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook: apply
* 11:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 11:34 jclark@cumin1004: START - Cookbook sre.hosts.reimage for host ml-serve1016.eqiad.wmnet with OS trixie
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 11:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/superset-next: apply
* 11:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 11:31 jclark@cumin1004: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 jclark@cumin1004: START - Cookbook sre.hosts.provision for host ml-serve1016.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 11:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 11:13 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2004.codfw.wmnet
* 11:13 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2003.codfw.wmnet
* 11:04 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' 'Martin Urbanec' --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:59 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2003.codfw.wmnet
* 10:54 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2002.codfw.wmnet
* 10:50 urbanecm@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:38 zabe: create wbc_entity_usage table in x1 for all wikidata client wikis # [[phab:T438499|T438499]]
* 10:36 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2002.codfw.wmnet
* 10:36 btullis@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:21 btullis@cumin1004: START - Cookbook sre.hosts.reboot-single for host dse-k8s-worker2001.codfw.wmnet
* 10:21 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki mediawikiwiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:20 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki metawiki 'Wikimedia Apps/Team/Android/Customizable Donation Reminder Experiment' 'Wikimedia Apps/Team/Customizable Donation Reminder/Android' Zabe --reason 'per request [[:phab:T438704{{!}}T438704]]'
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-eqiad@eqiad
* 10:17 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:16 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 10:12 jmm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: apply
* 10:11 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-eqiad@eqiad
* 10:09 jmm@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: apply
* 10:04 jmm@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: apply
* 10:02 jmm@deploy1003: helmfile [codfw] START helmfile.d/services/proton: apply
* 10:01 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 10:00 jmm@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: apply
* 09:59 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1063.eqiad.wmnet with OS trixie
* 09:59 jmm@deploy1003: helmfile [staging] START helmfile.d/services/proton: apply
* 09:56 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: wikikube-worker-codfw@codfw
* 09:55 jelto@cumin1004: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 jelto@cumin1004: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 09:54 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/liftwing-studio: apply
* 09:50 jelto@cumin1004: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: wikikube-worker-codfw@codfw
* 09:35 moritzm: installing chromium security updates
* 09:22 tappof: bump space for prometheus k8s-dse in eqiad
* 09:11 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:07 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 09:04 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:04 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:01 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]] (duration: 32m 54s)
* 09:01 filippo@cumin1004: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1063.eqiad.wmnet with reason: host reimage
* 08:58 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 08:45 filippo@cumin1004: START - Cookbook sre.hosts.reimage for host cloudvirt1063.eqiad.wmnet with OS trixie
* 08:29 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1341161{{!}}[Growth] Remove unused config variables (T392944)]]
* 08:15 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:05 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1063.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 08:04 filippo@cumin1004: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudvirt1063.eqiad.wmnet with reason: provision
* 08:01 XioNoX: restart gnmic on all netflow servers except 2005 and 1004 to pickup the new version - [[phab:T438291|T438291]]
* 07:59 XioNoX: install gnmic 0.49 on all netflow hosts - [[phab:T438291|T438291]]
* 07:57 XioNoX: add gnmic 0.49 to trixie-wikimedia - [[phab:T438291|T438291]]
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5a-codfw
* 07:53 ayounsi@cumin1004: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device fasw1-f5b-codfw
* 07:53 ayounsi@cumin1004: START - Cookbook sre.network.tls for network device fasw1-f5b-codfw
* 07:45 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:39 filippo@cumin1004: START - Cookbook sre.hosts.provision for host cloudvirt1077.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART
* 07:37 filippo@cumin1004: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudvirt1077.eqiad.wmnet
* 07:23 filippo@cumin1004: START - Cookbook sre.hosts.reboot-single for host cloudvirt1077.eqiad.wmnet
* 07:13 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:12 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:11 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 07:10 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 07:00 jmm@cumin2003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for krb1002.eqiad.wmnet: Renew puppet certificate - jmm@cumin2003
* 05:24 moritzm: upgrade docker-report on build2004 to 0.0.20 [[phab:T435314|T435314]]
* 05:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
== 2026-09-20 ==
* 20:08 dani@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 20:08 dani@deploy1003: helmfile [staging] DONE helmfile.d/services/miscweb: apply
* 20:07 dani@deploy1003: helmfile [staging] START helmfile.d/services/miscweb: apply
== 2026-09-19 ==
* 16:55 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:55 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:11 urbanecm: Attach SHB@commonswiki to the SUL account manually ([[phab:T438591|T438591]], see [[phab:T438591|T438591]]#12341750 for what I did exactly)
* 04:08 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 04:08 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 04:07 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
4ymakpbwlsscdvdd7tjcpe206nec8on
Release Engineering/SAL
0
17290
2461128
2461059
2026-09-26T20:11:33Z
Stashbot
7414
Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1345288
2461128
wikitext
text/x-wiki
=== 2026-09-26 ===
* 20:11 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1345288
=== 2026-09-25 ===
* 15:43 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 15:23 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 15:20 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.1-do.6 to 1.35.7-do.5 ([[phab:T439035|T439035]])
* 15:05 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.1-do.6 to 1.35.7-do.5 ([[phab:T439035|T439035]])
* 14:05 hashar: Tag Quibble 1.22.0 @ {{Gerrit|0d0bce00eedf44dd235703cb3eb1f60083d64752}} # [[phab:T437752|T437752]] [[phab:T426684|T426684]]
=== 2026-09-24 ===
* 17:02 dancy: Upgraded gitlab-cloud-runners (staging) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 16:01 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 15:37 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.1-do.6 to 1.35.7-do.5 ([[phab:T439035|T439035]])
* 15:30 dancy: Update gitlab-runner image to v19.2.6 ([[phab:T439037|T439037]])
=== 2026-09-23 ===
* 22:06 Southparkfan: decommission Beta Cluster PHP 8.3 hosts deployment-mediawiki13, deployment-mediawiki14, deployment-jobrunner05, deployment-mwmaint03 - [[phab:T435393|T435393]]
* 21:46 Southparkfan: cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - [[phab:T435393|T435393]]
* 20:37 hashar: gerrit: deleted https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344362 duplicate Change-Id {{Gerrit|I67dd7b3c879049a1c98586c365c297eac150b963}} of https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1329697
* 20:26 dancy: Upgrading istio to 1.30.5 on gitlab-cloud-runners (staging) ([[phab:T439035|T439035]])
* 20:13 hashar: gerrit: reindexing has been completed
* 19:38 hashar: gerrit: started reindexing for accounts, group and projects indices
* 19:36 hashar: gerrit: force online reindexing of the change index over ssh hitting gerrit1003 (switch over). `gerrit index start changes --force`
* 12:58 awight: add seanleong-wmde to deployment-prep
* 12:56 awight: manually run puppet agent
=== 2026-09-22 ===
* 22:43 Southparkfan: add wmgRedisLockPassword to PrivateSettings.php - [[phab:T436480|T436480]]
* 22:40 Southparkfan: create empty /etc/helmfile-defaults/mediawiki/release on -deploy04 and -deploy06 to unbreak Scap sync-masters - [[phab:T435393|T435393]]
* 21:22 Southparkfan: decommission deployment-cumin-3 - [[phab:T436470|T436470]]
* 21:20 Southparkfan: add deployment-deploy06 to scap::dsh::scap_masters and deployment_hosts - [[phab:T435393|T435393]]
* 20:53 Southparkfan: remove puppetserver cherry-pick for [[phab:T428052|T428052]], running newer version of HAProxy nowadays
* 20:08 Southparkfan: add deployment-jobrunner06 to 'jobrunner' dsh group - [[phab:T438656|T438656]]
* 18:17 dancy: Remove deployment-mwmaint03.deployment-prep.eqiad1.wikimedia.cloud from scap::dsh::groups.mediawiki-installation.host in deployment-prep project puppet config ([[phab:T435393|T435393]]) to stop wmf-beta-update-all spam.
* 08:53 hashar: Restarted CI Jenkins on contint1003 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343675 "jenkins: exit the JVM on OutOfMemoryError" # [[phab:T435791|T435791]]
* 07:21 awight: Purge opcache on deployment-jobrunner06.deployment-prep
=== 2026-09-21 ===
* 15:54 hashar: gerrit: on mediawiki/extensions/QuickSurveys deleted branch `master-backup` which was pointing at {{Gerrit|1ad54717c059bcd49093d902eab2c098b4efe91a}} (which is in `master`)
* 13:56 awight: Purging opcache on deployment-jobrunner06
* 11:37 hashar: Build docker-registry.wikimedia.org/releng/ajv:0.4.1-s5 (upgrade Node from 24.18.0 to 26.8.2) Used by PipelineLib # [[phab:T438085|T438085]]
=== 2026-09-18 ===
* 15:27 James_F: CI Node reverted to 24.
* 14:46 James_F: Zuul: [labs/tools/wdaudiolex-be] Install tox CI
* 14:01 James_F: Zuul: Remaining Node 24 -> 26 migrations
* 13:48 James_F: Zuul: Migrate MediaWiki-land independent CI Node from 24 to 26, for [[phab:T438085|T438085]]
=== 2026-09-17 ===
* 21:27 hashar: Re running `postmerge` for https://gerrit.wikimedia.org/r/c/mediawiki/services/wikifeeds/+/1342700 for [[phab:T438379|T438379]] # `zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/services/wikifeeds --change {{Gerrit|1342700}},1`
* 16:24 brett: delete deployment-cache, deployment-cache-text, and deployment-cache-upload prefixes - [[phab:T436468|T436468]]
* 15:59 hashar: gerrit: manually deleted old wmf branches from VisualEditor/VisualEditor # [[phab:T438373|T438373]]
* 15:09 hashar: gerrit: convert old REL branches on VisualEditor/VisualEditor ( [[phab:T380841|T380841]] [[phab:T428864|T428864]] [[phab:T428911|T428911]] ) using: for branch in REL1_25 REL1_26 REL1_27 REL1_28 REL1_29 REL1_30 REL1_31 REL1_32 REL1_33 REL1_34 REL1_35 REL1_36 REL1_37 REL1_38 REL1_39 REL1_40 REL1_41 REL1_42 REL1_44; do ./.tox/make-release/bin/python ./make-release/branch.py --delete --abandon --bundle ve $branch; done;
* 15:00 hashar: gerrit: converting REL1_44 branches to tags to formally EOL REL1_44 {{!}} [[phab:T428911|T428911]]
* 13:13 hashar: Updated plugins on the CI Jenkins
* 07:29 codders: rebuilt and restarted phpunit-results-cache server
=== 2026-09-16 ===
* 06:15 hashar: Updating node jobs from NodeJS 26.4.0 to 27.8.2 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1342056
=== 2026-09-15 ===
* 20:46 James_F: Docker: [quibble-bookworm, node24-test, node26-test] Drop jsduck etc. for [[phab:T363905|T363905]]
* 20:39 James_F: Zuul: […/OOJsUIAjaxLogin] Drop the JavaScript documentation job, for [[phab:T391706|T391706]]
* 19:09 Southparkfan: add deployment-cp-text09 and deployment-cp-upload09 to 'cache_hosts' hiera key - [[phab:T436468|T436468]]
* 18:33 James_F: Zuul: [labs/tools/ldap] Archive repository, for [[phab:T438076|T438076]]
* 18:32 James_F: Zuul: [integration/gear] Archive repository, for [[phab:T289512|T289512]] and [[phab:T438076|T438076]]
* 18:29 James_F: Zuul: [research/landing-page] Archive repository, for [[phab:T438076|T438076]]
* 18:28 James_F: Zuul: [cloud/toolforge/*] Archive two Toolforge repositories, for [[phab:T438076|T438076]]
* 18:26 James_F: Zuul: [node-rdkafka-factory, node-rdkafka-statsd] Archive repos, for [[phab:T366611|T366611]] and [[phab:T438076|T438076]].
* 18:24 James_F: Zuul: [operations/container/miscweb] Archive repository, for [[phab:T438076|T438076]]
* 18:22 James_F: Zuul: [mediawiki/services/recommendation-api] Archive service, for [[phab:T429123|T429123]] and [[phab:T438076|T438076]]
* 18:19 James_F: Zuul: [wikidata/query-builder, wikidata/query/gui] Archive repos, for [[phab:T438076|T438076]]
* 18:16 James_F: Zuul: [mediawiki/services/wikispeech/*] Archive the services, for [[phab:T344741|T344741]] and [[phab:T438076|T438076]]
* 10:38 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # corrupt Cypress cache? [[phab:T438002|T438002]]
* 10:00 hashar: Updating node based Jenkins jobs for https://gerrit.wikimedia.org/r/c/integration/config/+/1341713 {{!}} update npm jobs to drop debug logs from cache # [[phab:T437376|T437376]] [[phab:T426741|T426741]]
* 05:19 phedenskog: devel-stats upgraded datasette-dashboards from 0.7.1 to 0.8.0 for releng-data.wmcloud.org
=== 2026-09-14 ===
* 19:51 James_F: Zuul: [mediawiki/extensions/MobileApp] Add VisualEditor phan dep, for [[phab:T437736|T437736]]
* 19:49 Southparkfan: switch Beta Cluster to PHP 8.5 - [[phab:T435393|T435393]]
* 17:18 taavi: relaoding zuul to deploy https://gerrit.wikimedia.org/r/1337606
=== 2026-09-12 ===
* 18:23 Southparkfan: project-wide puppet hiera: remove puppetmaster::geoip::<nowiki>{</nowiki>fetch_private,use_proxy<nowiki>}</nowiki>, puppetdb_host, profile::puppetmaster::common::<nowiki>{</nowiki>command_broadcast,puppetdb_host,puppetdb_hosts<nowiki>}</nowiki>
* 18:14 Southparkfan: project-wide puppet hiera: remove role::puppetmaster::puppetdb::shared_buffers, obsoleted by profile::puppetdb::database::shared_buffers
* 18:09 Southparkfan: project-wide puppet hiera: remove profile::puppetdb::master pointing to former deployment-puppetdb02 (overridden in deployment-puppetdb prefix), profile::puppetdb::slaves (by default already empty) + profile::puppetdb::extra_authorized_hosts (variable does not exist, couldn't find it in historic commits either)
=== 2026-09-11 ===
* 16:29 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1339162
* 06:05 phedenskog: Updating Jenkins jobs from https://gerrit.wikimedia.org/r/c/integration/config/+/1338086 (quibble*, api-testing-*, mwext-phpunit-coverage*): run npm cache verify at most once a day per cache [[phab:T437377|T437377]]
=== 2026-09-10 ===
* 22:01 Southparkfan: decommission deployment-mwlog02 - [[phab:T436469|T436469]]
* 21:08 Southparkfan: masked and stopped stray udp2log.service on deployment-mwlog03, claimed 8420/udp which should have been assigned to udp_tee - [[phab:T436469|T436469]]
* 18:56 James_F: Zuul: [mediawiki/extensions/PersonalDashboard] Add ORES & Wikibase phan deps for [[phab:T436570|T436570]] and [[phab:T437491|T437491]]
* 18:55 James_F: Zuul: [mediawiki/extensions/ArticleGuidance] Add CommunityConfiguration for phan for [[phab:T437586|T437586]]
* 05:49 phedenskog: Updated quibble-for-mediawiki-core postgres and sqlite jobs to run only the phpunit-database stage [[phab:T437134|T437134]]
* 05:20 phedenskog: Updated 134 Quibble Jenkins jobs to drop the duplicated --reporting-url [[phab:T323750|T323750]]
=== 2026-09-09 ===
* 22:36 Southparkfan: point profile::rsyslog::udp_tee::destinations to both deployment-mwlog02 and deployment-mwlog03 8421/udp - [[phab:T436469|T436469]]
* 22:32 Southparkfan: set role::logging::mediawiki::udp2log::monitor: false to unbreak [[phab:T436469|T436469]]
* 16:53 Southparkfan: cpjobqueue: switch back to deployment-jobrunner05 (PHP 8.3) - [[phab:T435393|T435393]]
* 16:32 Southparkfan: adjust cpjobqueue config to temporarily send jobs to deployment-jobrunner06 (PHP 8.5) - [[phab:T435393|T435393]]
* 16:12 Southparkfan: decommission deployment-docker-mathoid02 - [[phab:T436465|T436465]]
* 16:09 thcipriani: deployed some anti-scraper mitigations on beta
* 15:40 komla@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.quota_increase (exit_code=0) by 30 server-groups ([[phab:T437117|T437117]])
* 15:40 komla@cloudcumin1001: START - Cookbook wmcs.openstack.quota_increase by 30 server-groups ([[phab:T437117|T437117]])
* 11:58 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.add_user_to_project (exit_code=0) for user 'denisse' in role 'member'
* 11:58 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.add_user_to_project for user 'denisse' in role 'member'
* 09:42 phedenskog: jjb: update quibble-with-gated-extensions-selenium-php83 to run npm install ahead of the browser tests [[phab:T436376|T436376]]
=== 2026-09-08 ===
* 19:36 Southparkfan: truncate Apache2 forensic log on deployment-mediawiki14, disk almost full
* 19:35 Southparkfan: set profile::mediawiki::httpd::enable_forensic_log to false, to avoid disk exhaustion during high-traffic load
* 16:21 thcipriani: hard reboot deployment-prep:deployment-cache-text08 horizon logs showing OOMs
* 15:42 hashar: Updating Quibble jobs to 1.21.0 # [[phab:T429715|T429715]] [[phab:T303270|T303270]] [[phab:T437134|T437134]] [[phab:T300727|T300727]]
* 15:06 hashar: Tag Quibble @ {{Gerrit|b692707e5213abf5e16894b1fb21cea41cd72c9e}} # [[phab:T429715|T429715]] [[phab:T303270|T303270]] [[phab:T437134|T437134]] [[phab:T300727|T300727]]
=== 2026-09-07 ===
* 19:31 Southparkfan: 19:31 UTC: switch back from PHP 8.5 to PHP 8.3 hosts - [[phab:T435393|T435393]]
* 19:06 Southparkfan: 18:05 UTC: pool mediawiki15 and mediawiki16 (PHP 8.5) as replacements for 13 and 14 (PHP 8.3), smoke test - [[phab:T435393|T435393]]
* 19:03 Southparkfan: banhammer lots of ranges to get Beta Cluster back online; not sure it was very effective, but we seem to be out of the woods
* 17:55 Southparkfan: no disk space left on deployment-mediawiki14, cleared logs in /var/log/apache2/forensic to unbreak
=== 2026-09-04 ===
* 15:52 hashar: integration: remove from Jenkins global config: NPM_CONFIG_AUDIT=false and NPM_CONFIG_FUND=false # [[phab:T437008|T437008]]
* 14:43 hashar: integration: set in Jenkins global config: NPM_CONFIG_AUDIT=false and NPM_CONFIG_FUND=false # [[phab:T437008|T437008]]
=== 2026-09-03 ===
* 23:29 Southparkfan: attach Cinder volume for /srv on deploy06, jobrunner06, mediawiki15/mediawiki16 - [[phab:T435393|T435393]]
* 11:26 hashar: Deleted coverage report for SimilarEditors ( /srv/doc/cover-extensions/SimilarEditors ), extension is being archived # [[phab:T436880|T436880]]
* 08:13 James_F: Zuul: [mediawiki/services/similar-users] Archive service, for [[phab:T368269|T368269]]
* 07:25 hashar: integration: granted sudo access to Phedenskog
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 05:15 hashar: Updating tox jobs to change default python from 3.9 to 3.11 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1334023 {{!}} [[phab:T436857|T436857]]
=== 2026-09-02 ===
* 22:21 Southparkfan: decommission deployment-webperf21 - [[phab:T436464|T436464]]
* 17:09 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1333891
* 15:37 hashar: jjb: update Quibble jobs to 1.20.0 # [[phab:T321617|T321617]] [[phab:T432966|T432966]] [[phab:T428642|T428642]] [[phab:T435975|T435975]]
* 15:14 hashar: Building Quibble 1.20.0 images
* 14:52 hashar: Tag Quibble 1.20.0 @ {{Gerrit|c6f13962c27685773fc1fb39d74189b202e0961d}} # [[phab:T321617|T321617]] [[phab:T432966|T432966]] [[phab:T428642|T428642]] [[phab:T435975|T435975]]
=== 2026-09-01 ===
* 21:45 Southparkfan: hiera: switch ATS routing for performance.beta.wmcloud.org from webperf21 to webperf31 - [[phab:T436464|T436464]]
* 20:55 Southparkfan: decommission deployment-webperf22 - [[phab:T436464|T436464]]
* 20:52 Southparkfan: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1333288 cherry-picked on project puppetserver - [[phab:T436464|T436464]]
* 19:59 Southparkfan: decommission deployment-webperf32 - [[phab:T436464|T436464]]
* 17:54 Southparkfan: switch webproxy for wikifeeds-beta to deployment-docker-wikifeeds01 - [[phab:T436462|T436462]]
* 17:47 Southparkfan: switch profile::restbase::citoid_uri and profile::restbase::cxserver_uri to resp. citoid03 and cxserver03, old VMs no longer exist - [[phab:T436619|T436619]]
* 17:43 Southparkfan: revoked Puppet certs for deployment-docker-cxserver02 and deployment-docker-citoid02 - [[phab:T436619|T436619]]
* 14:36 hashar: integration: deleted Cypress from codehealth job after the job learn to instruct Cypress & Puppeter to no more download binary blobs. `rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-codehealth-master-non-voting/Cypress` # [[phab:T427471|T427471]]
=== 2026-08-31 ===
* 22:17 andrewbogott: (log again, mentioned wrong task last time) add PHP 8.5 hosts to Scap dsh groups - [[phab:T435393|T435393]] (andrew retrying SPF's failed log)
* 21:21 Southparkfan: switched cache-text08 backend from mediawiki14 to mediawiki16, then switched back to mediawiki14 to match Puppet state - [[phab:T435393|T435393]]
* 21:11 Southparkfan: add PHP 8.5 hosts to Scap dsh groups - [[phab:T436462|T436462]]
* 18:38 Southparkfan: decommission deployment-wikifeeds02 - [[phab:T436462|T436462]]
* 13:21 jnuche: Updating development images on contint primary for [[phab:T435931|T435931]]
* 09:16 elukey: move cx-server and citoid-beta endpoints in deployment-prep to two new Trixie VMs, update their configs and Docker images and delete the old images.
* 04:18 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1332530
=== 2026-08-27 ===
* 15:06 Krinkle: Disable duplicate publishing noise from extension-IPReputation, EIPR, [[phab:T143162|T143162]]
=== 2026-08-26 ===
* 18:07 andrewbogott: resolving rebase conflicts in /srv/git/labs/private
* 17:24 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1329617
* 14:45 hashar: Updating Castor (0.4.3..0.4.5) on all Jenkins jobs {{!}} https://gerrit.wikimedia.org/r/1329577
* 07:11 hashar: integration: updated Quibble jobs to replace deprecated `--commands` option by 1..n `--command` option(s) {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1293711 {{!}} [[phab:T321617|T321617]]
=== 2026-08-25 ===
* 21:32 brett: Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08
* 21:23 brett: revert deployment-prep acme-chief switch: acme-chief active to deployment-acme-chief05 and passive to deployment-acme-chief06
* 21:03 brett: Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08
* 20:31 Southparkfan: [[phab:T401839|T401839]] - provisioned deployment-docker-wikifeeds01 (trixie) using Tofu (cloudvps-repos/deployment-prep/tofu-provisioning)
* 15:06 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/121 ([[phab:T435368|T435368]])
* 07:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.remove_instance (exit_code=0) for instance deployment-ircd03
* 07:35 taavi@cloudcumin1001: START - Cookbook wmcs.vps.remove_instance for instance deployment-ircd03
=== 2026-08-24 ===
* 08:29 hashar: integration: add firewall rule to ssh from contint1003/contint2003 IPv6 (IPv4 was already allowed) # [[phab:T435756|T435756]]
* 08:26 hashar: integration: removing firewall rule for ssh from contint1002/contint2002 IPv4 # [[phab:T418521|T418521]]
=== 2026-08-21 ===
* 13:07 hashar: integration: triggered Doxygen doc for PersonalDashboard extension using: `zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/extensions/PersonalDashboard --change {{Gerrit|1328044}},1` # [[phab:T435392|T435392]]
=== 2026-08-20 ===
* 23:26 bd808: deployment-mediawiki14: `systemctl stop php8.3-fpm; sleep 5m; systemctl start php8.3-fpm` -- maybe the bot storm will break if we give fast 500 responses for 5 minutes.
* 14:55 James_F: jforrester@integration-castor06:$ sudo rm -rf /srv/castor/castor-mw-ext-and-skins/master/wikilambda-catalyst-end-to-end # Clear stale Catalyst castor npm downloads.
* 13:35 James_F: Zuul: [mediawiki/extensions/WikiLambda] Re-enable Catalyst
* 08:54 hashar: deployment-prep: hard reboot deployment-cache-test-08 # [[phab:T435421|T435421]]
=== 2026-08-19 ===
* 14:48 hashar: integration: updated castor save job to have rsync emit statistics {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1321049 {{!}} [[phab:T432685|T432685]]
* 12:30 James_F: Updating development images on contint primary for “fundraising: Drop libc-client-dev from the bookworm PHP 8.2 image”
=== 2026-08-18 ===
* 17:32 dancy: Rebooting deployment-mediawiki14.deployment-prep.eqiad1.wikimedia.cloud for good measure
* 17:31 dancy: rm /var/log/apache2/*.gz on deployment-mediawiki14.deployment-prep.eqiad1.wikimedia.cloud to free up ~6GB.
* 09:24 hashar: zuul: restarted zuul-web
=== 2026-08-17 ===
* 21:09 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/115
* 21:01 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/118 ([[phab:T413817|T413817]], [[phab:T407430|T407430]])
* 18:08 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/113
* 16:30 James_F: Zuul: [mediawiki/extensions/WP25EasterEggs] Archive repository, for [[phab:T418134|T418134]]
=== 2026-08-14 ===
* 18:59 James_F: Zuul: Enforce CI for mediawiki-php-<nowiki>{</nowiki>excimer,luasandbox,wikidiff2<nowiki>}</nowiki> for [[phab:T425943|T425943]]
* 18:20 James_F: Docker: [php85] Migrate to Wikimedia-provide binary, cascaded, for [[phab:T433254|T433254]]
* 14:37 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1325900
=== 2026-08-13 ===
* 19:55 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/111 ([[phab:T401115|T401115]]) (redux for typo fix)
* 19:27 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/111 ([[phab:T401115|T401115]])
* 03:05 TimStarling: created Produnto tables on beta [[phab:T421436|T421436]]
=== 2026-08-12 ===
* 12:31 hashar: integration: on Castor: `sudo rm -fR /srv/castor/*/*/mwext-phpunit-coverage*/npm` # [[phab:T427922|T427922]]
* 11:59 hashar: integration: on Castor: `sudo rm -fR /srv/castor/*/*/*codehealth*/npm` # [[phab:T427822|T427822]]
=== 2026-08-07 ===
* 16:13 hashar: integration: deleted integration-agent-[[phab:T422258|T422258]] agent # [[phab:T422258|T422258]]
* 14:31 hashar: integration: updating Quibble jobs to Quibble 1.19.0 # [[phab:T432934|T432934]] [[phab:T432943|T432943]] [[phab:T427922|T427922]]
* 07:25 hashar: Tag Quibble 1.19.0 @ {{Gerrit|a8a84ed1c0adb34688ce22066ebde8a2d584a7d4}} # [[phab:T432934|T432934]] [[phab:T432943|T432943]] [[phab:T427922|T427922]]
=== 2026-08-06 ===
* 18:02 dancy: Restarted gitlab-webhooks ([[phab:T430410|T430410]]) (revert)
* 17:55 dancy: Restarted gitlab-webhooks ([[phab:T430410|T430410]])
=== 2026-08-05 ===
* 20:31 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1184176
* 18:39 James_F: Docker: [php85] Upgrade PHP to 8.5.9
* 15:58 James_F: Zuul: [mediawiki/extensions/WikiLambda] Disable Catalyst, corrupted npm cache
=== 2026-08-04 ===
* 19:06 dancy: Updating buildkitd to v0.32.2 on gitlab-cloud-runners (staging and production) ([[phab:T433985|T433985]])
* 16:48 James_F: Docker: [php83] Upgrade PHP to 8.3.33
=== 2026-08-03 ===
* 20:45 dancy: Updating buildkitd to v0.32.1 on gitlab-cloud-runners (staging and production) ([[phab:T433879|T433879]])
* 16:40 hashar: gerrit: added Vaughn Walters to integration group until he get added to the ciadmin LDAP group {{!}} [[phab:T433615|T433615]]
* 16:37 James_F: Zuul: Drop REL1_44 testing, EOL, for [[phab:T428911|T428911]]
=== 2026-07-31 ===
* 18:14 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/114
* 12:23 hashar: integration: sudo cumin --force -p 0 'name:docker' 'rm -fR /srv/jenkins/workspace/*pipeline*'
=== 2026-07-30 ===
* 17:28 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover/mediawiki-libs-node-cssjanus/ # [[phab:T424419|T424419]]
=== 2026-07-29 ===
* 21:33 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/109
* 18:33 dancy: Buildkit v0.32.0 deployed to gitlab-cloud-runners staging and production ([[phab:T433520|T433520]])
=== 2026-07-24 ===
* 17:59 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1316011
=== 2026-07-23 ===
* 15:06 dancy: Deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314124 ([[phab:T295351|T295351]])
=== 2026-07-22 ===
* 21:32 dancy: Moved /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/npm to /srv/castor-debug-[[phab:T20260722|T20260722]]-npm-torn-cacache/ on integration-castor06
* 18:41 dancy: Zuul dependencies upgraded and Zuul restarted.
* 18:25 dancy: Zuul is currently broken due to the Gerrit SSH key update. I'm investigating
* 17:24 brennen: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1314014/1 ([[phab:T432886|T432886]])
* 16:58 dancy: Restarting Gerrit ([[phab:T240266|T240266]]) (again)
* 15:59 dancy: Restarting Gerrit ([[phab:T240266|T240266]])
* 09:54 James_F: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure for {{Gerrit|1313321}}, per Lucas_WMDE.
=== 2026-07-21 ===
* 15:59 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1313220
* 07:08 hashar: integration: cleaned up Jenkins workspace on integration-agent-docker-1084
=== 2026-07-20 ===
* 19:36 dancy: Restarting jenkins on contint1003 to clear out hours-long stuck jobs
=== 2026-07-16 ===
* 23:07 mutante: gerrit1003/2002/2003: rm /srv/gerrit/.ssh/config and revert gerrit:1311577 which fixed [[phab:T398401|T398401]] but caused [[phab:T432413|T432413]] - replication is working again
* 18:51 dancy: Updated buildkit to v0.31.2 gitlab-cloud-runners (staging and production) ([[phab:T432360|T432360]])
* 18:37 dancy: Restarted Jenkins to unstick jobs.
* 18:33 dancy: Jenkins placed in shutdown mode in preparation for a restart
* 17:07 bd808: Hard reboot of deployment-cache-text08.deployment-prep.eqiad1.wikimedia.cloud via Horizon; console shows OOM ([[phab:T432374|T432374]])
* 16:40 brennen: contint1002: chown and chmod on /srv/zuul/git/mediawiki/extensions/CampaignEvents/.git per https://www.mediawiki.org/wiki/Continuous_integration/Zuul#Very_high_queue_of_merger:merge_functions
* 16:34 dancy: Restarted docker on contint2003 ([[phab:T432326|T432326]])
* 16:33 dancy: Restarted docker on contint1003 ([[phab:T432326|T432326]])
* 16:32 dancy: Restarted docker on contint1003
* 12:31 Krinkle: krinkle@contint1003:~$ sudo /usr/sbin/service jenkins restart
* 12:10 Krinkle: krinkle@contint1002: zuul restart
=== 2026-07-15 ===
* 18:58 James_F: dev-images: Re-build PHP images for latest point releases, for [[phab:T431099|T431099]]
* 09:13 James_F: Docker: [php83] Update PHP to 8.3.32 for [[phab:T431099|T431099]]
=== 2026-07-14 ===
* 11:53 James_F: Docker: [commit-message-validator] Update to v3.0.0, for [[phab:T431799|T431799]]
* 08:37 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud
=== 2026-07-13 ===
* 15:03 dancy: Updating gitlab-cloud-runners to v19.0.2
=== 2026-07-10 ===
* 11:56 James_F: Zuul: Disable all browser tests on release branches except Wikibase's, for [[phab:T430415|T430415]]
* 09:16 James_F: Docker: [commit-message-validator] Update to 2.3.0
=== 2026-07-09 ===
* 11:55 hashar: retriggering postmerge change for [[phab:T431582|T431582]]: zuul enqueue --trigger gerrit --pipeline postmerge --project machinelearning/liftwing/inference-services --change {{Gerrit|1308631}},3
=== 2026-07-08 ===
* 20:58 brennen: patchdemo: deployed https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/367 ([[phab:T427964|T427964]])
* 19:21 mutante: gerrit - replacing registerEmailPrivateKey in Gerrit config - this invalidates pending/outstanding email validation links for gerrit users - but does not affect active accounts or already verified email addresses
* 14:36 hashar: contint1003, contint2003: manually installed `docker-buildx` Debian package to validate https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308659 # [[phab:T431582|T431582]]
* 13:04 hashar: deployment-prep: git repack on /srv/mediawiki-staging/php-master
* 12:27 hashar: deployment-prep: on deployment server: clearing old branches for mediawiki/extensions and mediawiki/skins # [[phab:T428864|T428864]]
=== 2026-07-07 ===
* 19:41 mutante: contint1003/2003 - add jenkins-agent user to docker group; restart jenkins
* 17:17 hashar: gerrit: deleted /srv/gerrit/java_pid3571660.hprof
=== 2026-07-06 ===
* 20:06 dancy: Updated buildkitd to v0.31.1 in gitlab-cloud-runners ([[phab:T429988|T429988]])
* 08:11 James_F: Zuul: Add WikimediaAntiAbuse extension, for [[phab:T431023|T431023]]
=== 2026-07-03 ===
* 15:25 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add Elastica dep too
* 15:09 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CirrusSearch dep
* 11:22 James_F: Zuul: Make the in-mediawiki-tarball template real
* 10:06 James_F: Zuul: [mediawiki/extensions/TestKitchen] Don't drop from release branches
* 09:53 James_F: Docker: [quibble-coverage] Update phpunit-patch-coverage to 0.0.18, for [[phab:T423987|T423987]] and [[phab:T425807|T425807]]
=== 2026-07-02 ===
* 13:27 James_F: Docker: [composer-scratch] Upgrade composer to 2.10.2 and cascade, for [[phab:T428570|T428570]]
* 10:47 hashar: zuul1002: running Puppet agent to drop `wikimediacloud.org` from `no_proxy` {{!}} [[phab:T430479|T430479]]
=== 2026-07-01 ===
* 13:39 hashar: integration: added timestamping to operations-puppet-catalog-compiler and operations-puppet-catalog-compiler-puppet7-test jobs
* 02:45 hashar: gerrit: on gerrit2003 deleted /srv/gerrit/java_pid3520115.hprof (the JVM apparently died at some point, I assume due to heavy crawling)
=== 2026-06-30 ===
* 15:13 dancy: Rebooting deployment-mwlog02.deployment-prep to clear stuck udp2log processes
=== 2026-06-29 ===
* 09:32 hashar: gerrit: deleted repository phabricator/extensions/BurnDownCharts , created in July 2014, had no commit/changes
=== 2026-06-28 ===
* 15:30 hashar: Updated integration/zuul-jobs from upstream (c75fe6ef19c..fc4af6d4471), notably to remove `requestsexceptions` in `upload-logs-swift` role # [[phab:T430458|T430458]]
=== 2026-06-26 ===
* 13:59 Krinkle: [[phab:T429658|T429658]] krinkle@doc1004:/srv/doc/cover-extensions$ sudo -u doc-uploader rm -rf ShortUrl/
* 13:59 Krinkle: [[phab:T429658|T429658]] krinkle@doc2003:/srv/doc/cover-extensions$ sudo -u doc-uploader rm -rf ShortUrl/
=== 2026-06-25 ===
* 20:57 dancy: Restarting Jenkins to unstick builds
* 20:49 dancy: Investigating castor-save-workspace-cache clog
* 16:50 inflatador: add 60GB cinder vol to deployment-cirrussearch15 [[phab:T425585|T425585]]
* 14:25 inflatador: delete unused servers deployment-cirrussearch1[2-4] [[phab:T425585|T425585]]
* 10:07 James_F: Docker: Bump Node 24 / Node 26 to new releases
=== 2026-06-24 ===
* 17:26 dancy: Set `profile::puppetserver::autosign: /usr/local/sbin/validatecloudvpsfqdn.py` in hiera config for deployment-puppetserver prefix ([[phab:T429413|T429413]])
* 15:12 dancy: sudo systemctl restart php8.3-fpm on deployment-jobrunner05 (Attempting to resolve logspam)
=== 2026-06-23 ===
* 14:31 brennen: deploying patchdemo for https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/361
=== 2026-06-22 ===
* 23:19 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/mediawiki-core/master/mediawiki-node24/ #[[phab:T429824|T429824]]
* 23:07 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ #[[phab:T429824|T429824]] (again)
* 22:35 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ #[[phab:T429824|T429824]]
* 19:30 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-with-gated-extensions-vendor-mysql-php83 #[[phab:T429824|T429824]]
=== 2026-06-19 ===
* 20:13 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1304632
* 09:39 Lucas_WMDE: deployment-deploy04: used createAndPromote to restore User:Lucas Werkmeister (WMDE) to bureaucrat (dewiki, enwiki, metawiki, wikidatawiki) and wikidata-staff (wikidatawiki) after groups were very unhelpfully removed, presumably due to 2FA enforcement, without apparent warning or announcement of any kind
* 07:07 dcausse: restarted php-fpm on deployment-jobrunner05 (inconsistent php state: MediaWiki\Extension\PageAssessments\HookHandler\ParserHooks::__construct(): Argument #3 ($config) must be of type MediaWiki\Config\Config, MediaWiki\Page\WikiPageFactory given)
* 06:13 hashar: Upgrading Quibble jobs to 1.18.2 (retry git connections on reset) https://gerrit.wikimedia.org/r/c/integration/config/+/1304181 # [[phab:T420865|T420865]]
=== 2026-06-18 ===
* 19:52 dancy: Building quibble 1.18.2 images on contint primary ([[phab:T420865|T420865]])
* 18:50 dancy: Tag Quibble 1.18.2 @ {{Gerrit|152b497d317eaafc3cec334d5ce7e549697a2980}} # [[phab:T420865|T420865]]
* 15:15 dancy: Deleting deployment-db11 and deployment-db14 ([[phab:T428910|T428910]])
* 07:23 dcausse: reindexing all wikis to opensearch2 ([[phab:T425585|T425585]], [[phab:T427196|T427196]])
=== 2026-06-17 ===
* 23:16 dduvall: restarted zuul to clear up 5 hrs of stuck queues
* 23:03 mutante: re-enabled puppet on contint1003 - triple checked puppet does NOT start jenkins anymore. BOTH masked AND stopped while running untouched on contint1002. after: gerrit:1303578 {{!}} ([[phab:T418521|T418521]]) ([[phab:T428791|T428791]])
* 18:51 dduvall: performed stop/start of jenkins service on contint1002 following failed safeRestart
* 18:45 dduvall: restarting jenkins due to stuck zuul queues
* 14:56 inflatador: mwscript /srv/mediawiki-staging/php-master/extensions/CirrusSearch/maintenance/ForceSearchIndex.php --wiki=enwikibooks
=== 2026-06-16 ===
* 16:00 dancy: sudo keyholder arm on deployment-deploy04.deployment-prep
* 15:48 dancy: Resizing deployment-deploy04.deployment-prep from g4.cores4.ram8.disk20 to g4.cores8.ram16.disk20 ([[phab:T429364|T429364]])
* 15:31 dancy: Turning off deployment-db11 and deployment-db14
=== 2026-06-15 ===
* 23:40 dancy: systemctl restart php8.3-fpm on deployment-jobrunner05 to reload beta db configuration ([[phab:T428930|T428930]])
* 21:00 dancy: deployment-db15 promoted to master ([[phab:T428930|T428930]])
* 20:53 dancy: deployment-db11.deployment-prep going read-only
* 20:48 dancy: deployment-db11.deployment-prep will be going read-only soon while db15 is being promoted to primary.
* 20:26 dancy: Added deployment-db16 ([[phab:T429245|T429245]])
* 18:26 dancy: Rebooting deployment-jobrunner05
=== 2026-06-12 ===
* 21:45 dancy: Unstuck wmf-beta-update-all service on deployment-deploy04.deployment-prep (sudo systemctl stop wmf-beta-update-all)
* 18:00 thcipriani: unmasking jenkins on contint1002 and restarting
* 17:49 thcipriani: attempting to cancel castor-save-workspace-cache {{Gerrit|6710545}}
* 15:19 James_F: Docker: [php83] Re-platform to Debian Bookworm, for [[phab:T383337|T383337]]
* 15:07 dancy: deployment-db15 configured as a replica of deployment-db11 ([[phab:T428930|T428930]])
* 10:21 Krinkle: `krinkle@<nowiki>{</nowiki>doc1004,doc2003<nowiki>}</nowiki>:/srv/doc/mediawiki-core$ sudo -u doc-uploader rm -rf list/` - remove doc build for git-tag test.
=== 2026-06-11 ===
* 14:46 hashar: for minor in $(seq 21 42); do ./.tox/make-release/bin/python -u ./make-release/branch.py --delete --abandon --bundle '*' "REL1_$minor"; done;
* 14:46 hashar: On all MediaWiki repos, converting old release branches up to REL1_42 included to tags. Last time I missed non wmf repo # [[phab:T380841|T380841]] {{!}} [[phab:T428864|T428864]]
* 09:45 hashar: Converted mediawiki/core branches REL1_39, REL1_40, REL1_41, REL1_42 to tags # [[phab:T428864|T428864]]
* 09:18 hashar: Converting REL1_42 branches to tags # [[phab:T428864|T428864]]
* 09:18 hashar: Converting REL1_41 branches to tags # [[phab:T428864|T428864]]
* 09:05 hashar: Converting REL1_40 branches to tags # [[phab:T428864|T428864]]
* 08:56 hashar: Converting REL1_39 branches to tags # [[phab:T428864|T428864]]
* 08:40 hashar: gerrit: deleted mediawiki/core branch "development" that pointed to {{Gerrit|7f622781cc31053b121f6f4ddbff506cba10d38e}} (which is contained by master). Had probably been created by a direct push.
=== 2026-06-10 ===
* 16:06 James_F: Docker: Provide quibble-bookworm, for [[phab:T362705|T362705]]
=== 2026-06-04 ===
* 23:39 jeena: Updating development images on contint primary for [[phab:T424691|T424691]]
* 12:30 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc # fix failure seen in mwext-node24-rundoc 4812
* 08:35 hashar: Built Docker images `docker-registry.wikimedia.org/releng/java21:0.1` and `docker-registry.wikimedia.org/releng/maven-java21:0.1` # [[phab:T412978|T412978]]
=== 2026-06-03 ===
* 14:19 jnuche: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/107
* 12:54 James_F: Zuul: Add Rae 5e as a trusted user
* 08:26 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1296559 "inference-services: Add LLM generated editing suggestions CI/CD pipelines." # [[phab:T427794|T427794]]
=== 2026-06-02 ===
* 23:23 thcipriani: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/1296695
* 23:10 thcipriani: tag quibble 1.18.1 @ {{Gerrit|4b7959553c095d811426f394165c69ecc13a44eb}}
* 18:13 brennen: devtools phab/phorge: deployed work/2026-06-01-merge-phorge to https://phabricator.wmcloud.org/ for testing ([[phab:T410849|T410849]])
* 16:05 hashar: integration: upgraded pypy from 7.3.11 (py3.9) to 7.3.20 (py3.11) # [[phab:T423607|T423607]]
* 15:49 jnuche: Updating buildkitd to v0.30.0 in gitlab-cloud-runners ([[phab:T426212|T426212]])
* 15:24 hashar: Building docker images for https://gerrit.wikimedia.org/r/c/integration/config/+/1295865 # [[phab:T423607|T423607]]
* 14:57 jnuche: Jenkins/Zuul is back
* 14:38 jnuche: restarting Jenkins
* 14:28 jnuche: bring back castor node, that didn't help
* 14:23 jnuche: trying to reconnect castor node, see if that helps somehow
* 14:12 jnuche: option to "Enable Gearman" times out. Can't re-enable from UI. Gearman plugin logs are empty. Neat
* 14:05 jnuche: trying to reconnect Gearman
=== 2026-06-01 ===
* 21:34 jeena: Updating development images on contint primary for [[phab:T424663|T424663]]
* 10:58 hashar: gerrit: flushed `ldap_usernames` cache in case a missing account ended up being cached there # [[phab:T427792|T427792]]
=== 2026-05-28 ===
* 16:48 hashar: castor: nuked SonarQube cache: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-codehealth-master-non-voting/sonar/ # [[phab:T427471|T427471]]
* 07:05 hashar: integration: delete deployment-deploy04 agent from Jenkins controller. Batch job has been migrated to a systemd driven script # [[phab:T256168|T256168]]
* 06:29 hashar: integration: delete integration-castor05 agent from Jenkins, replaced by integration-castor06 # [[phab:T421114|T421114]]
=== 2026-05-27 ===
* 21:07 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1294425
* 09:49 hashar: Updated helm-lint Jenkins job to use releng/helm-linter:0.8.0 image # [[phab:T424824|T424824]]
* 08:25 codders: integration-castor06: sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/
=== 2026-05-26 ===
* 15:26 James_F: Zuul: Add PHP 8.5 to a few missing PHP pipelines, oops
* 14:46 hashar: Updating Quibble Jenkins jobs to stop Supervisord from spawning Memcached # [[phab:T397810|T397810]]
* 14:07 hashar: Updated integration-quibble-* jobs in order to validate running tests with Supervisord not managing memcached # [[phab:T397810|T397810]]
* 13:15 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-docs-publish # fix failure seen in mwext-node24-docs-publish 981, 985, 987
* 09:05 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1292620 (Introduce Phan composer job - [[phab:T231966|T231966]])
=== 2026-05-21 ===
* 16:33 hashar: Reloaded Zuul to enable Node24 CI job for `labs/tools/wdaudiolex-fe` # [[phab:T426366|T426366]]
=== 2026-05-19 ===
* 18:07 James_F: Zuul: [mediawiki/libs/ZestJQ] Add basic PHP and Node CI
=== 2026-05-18 ===
* 21:12 James_F: Zuul: [operations/software/gerrit] Add Node26 as experimental
* 21:10 mutante: gerrit-replica.wikimedia.org, gerrit-spare.wikimedia.org - rebooting backends
* 20:57 James_F: Zuul: [integration/docroot] Test in PHP 8.3+, dropping 8.2
* 20:56 James_F: Zuul: [analytics/wmde/scripts] Test in PHP 8.3+, dropping 8.
* 20:18 James_F: Zuul: [wikimedia/fundraising/dash] Replace Node 20 testing with Node 24
* 20:18 James_F: Zuul: Migrate various labs things to Node 24
* 20:02 James_F: Docker: [ajv, sonar-scanner] Migrate to Node 24
* 19:58 James_F: Zuul: Migrate various production/CI things to Node 24
* 18:17 mutante: releases.wikimedia.org - rebooting backends
* 18:14 mutante: rebooting production gitlab-runners
* 18:12 dancy: gitlab-cloud-runners have been revived.
* 18:11 James_F: Zuul: [design/codex] Switch CI to Node 24
* 15:52 dancy: gitlab-cloud-runners are in a broken state. I'm investigating
* 14:27 hashar: Upgrading Quibble jobs to 1.18.0
* 09:29 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure seen in mediawiki-node24 22260
=== 2026-05-15 ===
* 18:36 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 18:24 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 18:11 dancy: Upgraded gitlab-cloud-runners (staging) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 17:59 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 13:02 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1287828 [[phab:T426392|T426392]]
=== 2026-05-13 ===
* 12:42 James_F: Zuul: [mediawiki/extensions/Springboard] Add AdminLinks Phan dependency
* 12:42 James_F: Zuul: [mediawiki/extensions/ChatBot] Add dependencies on VisualEditor and BlueSpiceFoundation
* 12:42 James_F: Zuul: [mediawiki/extensions/ChatIntegration] Add dependency on VisualEditor
* 12:37 James_F: Zuul: [mediawiki/extensions/WikiLambda] Drop AF and SB deps down to phan-only, for [[phab:T423180|T423180]]
=== 2026-05-12 ===
* 20:57 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/104 ([[phab:T424774|T424774]])
* 18:08 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add AF and SB deps for [[phab:T423180|T423180]]
* 14:18 atsukoito: PrivateSettings: empty $wgOpensearchCredentials for opensearch-on-k8s synced to deploy04 by Reedy
* 13:04 atsukoito: PrivateSettings: credentials for opensearch-on-k8s ttmserver-test
* 11:50 James_F: Zuul: [machinelearning/liftwing/inference-services] Add qwen36 llm model CI/CD pipelines, for [[phab:T425680|T425680]]
* 11:46 James_F: Zuul: Add experimental php-pie-build* jobs to other PHP extensions, for [[phab:T425943|T425943]]
* 11:37 James_F: Zuul: [mediawiki/php/wikidiff2] Add experimental php-pie-build* jobs, for [[phab:T425943|T425943]]
* 10:05 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-with-Wikibase-extensions-browser-tests-only-vendor-php83 # fix failure seen in quibble-with-Wikibase-extensions-browser-tests-only-vendor-php83 7817
* 08:44 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # broken Cypress cache? hopefully fix failure seen in quibble-vendor-mysql-php83-selenium 51633
=== 2026-05-11 ===
* 18:28 James_F: Docker: Add changes to php-compile images for PIE, for [[phab:T425943|T425943]]
* 16:06 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # broken Cypress cache? hopefully fix failure seen in quibble-vendor-mysql-php83-selenium 51439 and 51452
=== 2026-05-09 ===
* 20:46 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1285498
=== 2026-05-07 ===
* 22:53 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/105
=== 2026-05-06 ===
* 18:13 bd808: Unblock 88.165.192.0/19
* 18:03 bd808: Unblock 94.208.0.0/14
* 17:56 bd808: Unblock 84.226.0.0/16
* 17:41 bd808: Unblock 94.34.0.0/16
* 17:35 bd808: Unblock 109.134.0.0/16
=== 2026-05-05 ===
* 21:20 James_F: Zuul: Provide Node 26 experimental jobs everywhere needed
* 21:04 James_F: Docker: Provide initial Node 26 images
* 19:01 James_F: Zuul: [mediawiki/extensions/PageAssessments] Add Scribunto dependency, for [[phab:T396135|T396135]]
* 14:58 dancy: rm /var/log/<nowiki>{</nowiki>user.log.1,syslog.1,messages.1<nowiki>}</nowiki> on deployment-eventgate-4.deployment- prep ([[phab:T425429|T425429]])
=== 2026-05-04 ===
* 15:19 dancy: Upgrading gitlab cloud runners (prod) from 1.35.1-do.3 to 1.35.1-do.5
* 14:51 dancy: Upgrading gitlab cloud runners (staging) from 1.35.1-do.3 to 1.35.1-do.5
* 10:40 James_F: Zuul: Provide non-voting PHP 8.4/8.5 Quibble jobs for bluespice template
=== 2026-05-02 ===
* 20:49 James_F: Zuul: [mediawiki/core] Enforce PHP 8.4 & 8.5 on release branches, all pass
* 19:27 James_F: Zuul: Provide non-voting PHP 8.4/8.5 Quibble jobs for MW release branches
* 19:19 James_F: Zuul: [mediawiki/extensions/BlogPage] Add dependencies
* 16:48 James_F: Hard-restarting Zuul to clear the huge number of i18n updates being re-submitted.
* 15:48 James_F: Zuul: [wikimedia-cz/*] Test in PHP 8.3+, dropping 8.2
* 14:02 TheresNoTime: Add bvibber to deployment-prep project
* 09:08 James_F: Docker: [quibble-*] Add php-luasandbox so we can test both modes in Scribunto
=== 2026-05-01 ===
* 15:42 James_F: Zuul: [wikimedia/lucene-explain-parser] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: Zuul: [wikimedia/textcat] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: Zuul: [mediawiki/tools/ParseWiki] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: zuul: Add ToprakM to CI allowlist
* 15:19 James_F: Zuul: [translatewiki] Test in PHP 8.3+, dropping 8.2
* 15:10 James_F: Zuul: [mediawiki/extensions/WikiEditor] Add TestKitchen as a dependency, for [[phab:T425076|T425076]]
* 12:40 James_F: Zuul: [mediawiki/tools/code-utils] Test in PHP 8.3+, dropping 8.2
* 08:02 James_F: Zuul: Update xtex's e-mail in the allowlist
* 07:37 James_F: Zuul: Switch release branches' selenium jobs to PHP 8.3
* 07:33 James_F: Zuul: Test Wikimedia production libraries in PHP 8.3+, dropping 8.2
=== 2026-04-30 ===
* 21:36 brennen: gitlab-webhooks: building & restarting to deploy https://gitlab.wikimedia.org/repos/releng/gitlab-webhooks/-/merge_requests/40
* 20:26 James_F: Zuul: [mediawiki/tools/api-testing] Make PHP 8.5 CI voting
* 20:16 James_F: jforrester@doc1004:~$ # sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/WebAuthn/ # [[phab:T415832|T415832]]
* 20:14 James_F: Zuul: [mediawiki/extensions/WebAuthn] Archive, for [[phab:T415832|T415832]] / [[phab:T303495|T303495]]
* 17:16 brennen: wikibugs: most maintainers at hackathon, so go release-engineering added as a maintainer while looking to debug error at https://gitlab.wikimedia.org/toolforge-repos/wikibugs2/-/jobs/810904
* 15:19 mutante: upgrading zuul to 14.2.0-1 on "new zuul" machines ([[phab:T424879|T424879]])
=== 2026-04-29 ===
* 15:49 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add ConfirmEdit dependency, for [[phab:T424597|T424597]]
* 15:36 James_F: Zuul: Drop experimental node22 jobs, never used in practice
* 15:28 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1279392, https://gerrit.wikimedia.org/r/1279397
=== 2026-04-28 ===
* 18:11 bd808: Unblock 86.0.0.0/16
* 17:41 bd808: Unblock 79.192.0.0/10
* 17:07 James_F: Zuul: [mediawiki/tools/phpunit-patch-coverage] Drop PHP 8.2 testing
* 17:07 James_F: Zuul: [mediawiki/tools/minus-x] Drop PHP 8.2 testing
* 17:07 James_F: Zuul: [mediawiki/tools/codesniffer] Drop PHP 8.2 testing
* 16:32 James_F: Zuul: [mediawiki/services/jobrunner] Drop PHP 8.2 testing
* 13:34 James_F: Zuul: [mediawiki/tools/phan] Drop PHP 8.2 testing
* 13:34 James_F: Zuul: [oojs/ui] Drop PHP 8.2 testing
* 13:14 James_F: Zuul: [mediawiki/tools/phan/SecurityCheckPlugin] Drop PHP 8.2 CI
* 10:40 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud to fix failure seen in mwext-node24-rundoc #1717
* 00:03 bd808: Increase parallelism for wmf-beta-update-databases.py ([[phab:T256168|T256168]])
=== 2026-04-27 ===
* 22:11 bd808: Beta Cluster MediaWiki update logs now available via https://beta-update.wmcloud.org/ ([[phab:T256168|T256168]])
* 21:57 bd808: Add web security group to deployment-deploy04 ([[phab:T256168|T256168]])
* 20:45 James_F: Zuul: Restrict mw*-codehealth-patch jobs to master only, for [[phab:T424573|T424573]]
* 17:16 James_F: Docker: [mediawiki-phan-taint-check-demo] Re-platform to Trixie and so PHP 8.4
* 15:53 James_F: Zuul: [mediawiki/extensions/ReportIncident] Add TestKitchen phan dependency, for [[phab:T424220|T424220]]
* 14:32 James_F: Zuul: Drop PHP 8.2 enforcement from MediaWiki things for master and REL1_46 for [[phab:T358667|T358667]]
* 12:38 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-docs-publish # fix failure seen in mwext-node24-docs-publish 383
* 09:18 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover/mediawiki-libs-node-cssjanus/ # [[phab:T424419|T424419]]
=== 2026-04-26 ===
* 20:49 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1276777
=== 2026-04-24 ===
* 22:48 dduvall: merged zuul3 branch of integration/config into master and pushed (in preparation for https://gerrit.wikimedia.org/r/c/operations/puppet/+/1277198)
* 12:27 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1276428
=== 2026-04-23 ===
* 23:57 bd808: Set `profile::beta::autoupdater::run_updater: true` for deployment-deploy04 via Horizon ([[phab:T256168|T256168]])
* 22:58 bd808: bd808@deployment-deploy04 `sudo -u jenkins-deploy /usr/local/bin/wmf-beta-update-all`
* 22:36 bd808: bd808@deployment-deploy04 `sudo -u mwdeploy /usr/local/bin/wmf-beta-update-all`
* 22:16 bd808: Disabled https://integration.wikimedia.org/ci/view/Beta/job/beta-update-databases-eqiad so that replacement script can be tested ([[phab:T256168|T256168]])
* 22:12 bd808: Disabled https://integration.wikimedia.org/ci/job/beta-code-update-eqiad so that replacement script can be tested ([[phab:T256168|T256168]])
* 22:02 bd808: Cherry-picked {{gerrit|1276813}} to deployment-puppetserver-1 ([[phab:T256168|T256168]])
* 20:11 James_F: Zuul: [wikibase/*] Replace CI testing in Node 20 with Node 24
* 20:11 James_F: Zuul: [wikidata/query/*] Replace CI testing in Node 20 with Node 24
* 20:11 James_F: Zuul: [analytics/*] Replace CI testing in Node 20 with Node 24
* 20:10 James_F: Zuul: [mediawiki/tools/*] Replace CI testing in Node 20 with Node 24
* 20:06 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.34.5-do.3 to 1.35.1-do.3 ([[phab:T423726|T423726]])
* 19:55 James_F: Zuul: [jquery-client] Replace CI testing in Node 20 with Node 24
* 19:51 James_F: Zuul: [wikipeg] Drop testing in Node 20 and Node 22
* 19:47 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.34.5-do.3 to 1.35.1-do.3 ([[phab:T423726|T423726]])
* 19:37 James_F: Zuul: [oojs/ui] Drop CI testing in Node 20 and Node 22
* 19:37 James_F: Zuul: [oojs/js] Drop CI testing in Node 20 and Node 22
* 19:37 James_F: Zuul: [unicodejs] Replace CI testing in Node 20 with Node 24
* 19:36 James_F: Zuul: [wikimedia/portals] Drop CI testing in Node 20 and Node 22
* 18:57 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.33.9-do.3 to 1.34.5-do.3 ([[phab:T423726|T423726]])
* 18:39 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.33.9-do.3 to 1.34.5-do.3 ([[phab:T423726|T423726]])
* 18:18 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.33.9-do.2 to 1.33.9-do.3 ([[phab:T423726|T423726]])
* 17:58 James_F: Zuul: [mediawiki/extensions/OAuth] Add dependency on CentralAuth, for [[phab:T415281|T415281]]
* 17:56 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.32.13-do.2 to 1.33.9-do.3 ([[phab:T423726|T423726]])
* 16:35 James_F: Zuul: Enforce PHP 8.5 CI for MW things in master (and REL1_46), for [[phab:T411814|T411814]]
* 16:19 James_F: Zuul: [mediawiki/services/parsoid] Enable PHP 8.5 CI
* 15:47 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add AntiSpoof dependency, for [[phab:T420548|T420548]]
* 14:20 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure seen in mediawiki-node24 8385 and 8405
* 12:56 James_F: Zuul: [mediawiki/extensions/GrowthExperiments] Add CentralNotice dependency, for [[phab:T422082|T422082]]
=== 2026-04-22 ===
* 00:07 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add MF dependency, for [[phab:T424113|T424113]]
=== 2026-04-21 ===
* 23:26 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CommunityConfiguration dep too, for [[phab:T394410|T394410]]
* 23:17 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add standalone test jobs, for [[phab:T422031|T422031]]
* 20:47 inflatador: updating cirrussearch hosts to Trixie/OpenSearch 2 [[phab:T421763|T421763]]
* 20:38 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CommunityConfiguration phan dep, for [[phab:T394410|T394410]]
* 20:17 bd808: Running tofu for [[phab:T421244|T421244]]
* 18:00 James_F: Zuul: [mediawiki/extensions/WatchAnalytics] Add ApprovedRevs Phan dependency
* 16:35 bd808: Unblock 79.116.0.0/16
* 13:34 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add TestKitchen phan dep, for [[phab:T415254|T415254]]
* 13:27 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add CentralAuth dependency, for [[phab:T420548|T420548]]
=== 2026-04-20 ===
* 23:56 bd808: Unblock 76.157.0.0/16
* 18:28 dancy: Upgrading gitlab cloud runners (staging) to 1.33.9-do.2 ([[phab:T423726|T423726]])
* 18:28 dancy: Upgrading gitlab cloud runners (staging) ([[phab:T423726|T423726]])
* 18:19 James_F: jjb: All 486 (!) jobs now updated for [[phab:T423622|T423622]]
* 18:18 bd808: Unblock 113.128.0.0/15
* 15:03 James_F: Docker: Bump ci-bullseye/-bookworm/-trixie for mirrors.wm.org removal, [[phab:T423622|T423622]]
=== 2026-04-19 ===
* 19:53 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1272752
=== 2026-04-17 ===
* 21:07 thcipriani: marking integration-agent-1080 offline for experimentation
* 19:30 thcipriani: reconfiguring castor-save-workspace-cache with https://gerrit.wikimedia.org/r/1273935
* 17:47 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.32.10-do.1 to 1.32.13-do.2 ([[phab:T423726|T423726]])
* 16:49 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.32.10-do.1 to 1.32.13-do.2 ([[phab:T423726|T423726]])
=== 2026-04-16 ===
* 20:49 dduvall: creating integration/zuul-jobs repo to serve as a mirror of opendev.org/zuul/zuul-jobs ([[phab:T406384|T406384]])
* 13:38 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1272711 [[phab:T423568|T423568]]
* 11:07 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud
=== 2026-04-15 ===
* 20:05 James_F: Zuul: Configure REL1_46 CI, for [[phab:T423257|T423257]]
* 17:44 bd808: Unblock 176.0.0.0/13
* 17:39 bd808: Unblock 46.128.0.0/16
* 17:32 bd808: Unblock 176.86.0.0/16
* 16:39 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/commit/127d783b2176ac60b646a5fa4f1b1a872ca66340
* 15:33 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/100
* 01:02 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/99
=== 2026-04-14 ===
* 20:42 James_F: Docker: [composer-scratch] Upgrade composer to 2.9.7 and cascade
* 16:35 bd808: Unblock 88.112.0.0/14
* 00:48 bd808: Unblock 24.6.0.0/16
* 00:42 bd808: Unblock 152.231.48.0/20
=== 2026-04-13 ===
* 22:00 James_F: Zuul: [mediawiki/vendor] Drop accidental Wikibase browser tests on branches
* 20:28 James_F: Zuul: [mediawiki/extensions/Chart] Drop Doxygen publish job, not used
* 14:42 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add FlaggedRevs dep, for [[phab:T421011|T421011]]
=== 2026-04-12 ===
* 18:21 James_F: jforrester@contint1002:~$ sudo /usr/sbin/service zuul restart && tail -f -n100 /var/log/zuul/zuul.log # [[phab:T423027|T423027]]
=== 2026-04-10 ===
* 23:22 James_F: jforrester@contint1002:~$ zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/extensions/ReadingLists --change {{Gerrit|1269498}},2 # [[phab:T422976|T422976]]
* 23:20 James_F: Zuul: [mediawiki/extensions/ReadingLists] Publish JS coverage, for [[phab:T422976|T422976]]
* 23:13 James_F: Zuul: Migrate a few straggler Node 20 MediaWiki things to Node 24
* 23:01 James_F: Zuul: Move all MediaWiki things from mediawiki-node20 to mediawiki-node24
* 21:59 James_F: Docker: Bump Node base images to March releases and cascade; Upgrade Quibble images from Node 20 to Node 24
* 10:24 hashar: Updating all Quibble jobs to 1.17.1
* 10:22 hashar: Updated PostgreSQL jobs to Quibble 1.17.1 # [[phab:T422110|T422110]]
* 10:22 hashar: Updated apitesting job to Quibble 1.17.1 # [[phab:T422843|T422843]] [[phab:T418743|T418743]]
* 09:51 hashar: Tag Quibble 1.17.1 @ {{Gerrit|0a1ab3b7c3dfee36c9bc2e9b049957d94e190e85}}
=== 2026-04-09 ===
* 15:13 hashar: Rolling back Quibble jobs to 1.16.0 (api-testing stage fails due to missing npm install step`
* 14:58 hashar: Upgrading Quibble jobs to 1.17.0
* 14:23 hashar: Tagged Quibble 1.17.0 @ {{Gerrit|864381c6b63bdbcd8c74a3162c406fffcaaf8694}}
* 07:48 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1268559 "Zuul: use standalone jobs for GrowthExperiments Cypress tests" {{!}} [[phab:T417412|T417412]]
=== 2026-04-08 ===
* 22:19 dancy: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1269068
* 22:01 bd808: Unblock 95.216.12.170/32 ([[phab:T422751|T422751]])
* 19:26 brennen: gitlab-webhooks: building & deploying https://gitlab.wikimedia.org/repos/releng/gitlab-webhooks/-/merge_requests/37 - hitting some build tooling stuff, trying a fix per instructions in the error log
* 17:54 bd808: Unblock 167.56.0.0/13 ([[phab:T422721|T422721]])
* 06:31 hashar: Deleted integration-agent-castor05 Bullseye instance, replaced by integration-agent-castor06 which is on Bookworm # [[phab:T421114|T421114]]
* 06:24 hashar: Deleted integration-agent-qemu-1003 Bullseye image, replaced by integration-agent-qemu-1004 which is on Bookworm # [[phab:T422488|T422488]]
=== 2026-04-07 ===
* 22:25 dduvall: adding new pipelinelib labels to ci nodes ([[phab:T422234|T422234]])
* 20:05 hashar: Triggered a build of https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/
* 17:06 dduvall: added `Docker` label to `contint` jenkins nodes ([[phab:T422507|T422507]])
* 17:05 dduvall: restored missing `pipelinelib` labels on `integration-agent-docker-` CI hosts ([[phab:T422507|T422507]])
* 16:53 bd808: Unblock 73.0.0.0/8 ([[phab:T422498|T422498]])
* 12:36 hashar: jjb: use $CASTOR_HOST for Quibble success cache. https://gerrit.wikimedia.org/r/1268545 {{!}} This causes the Quibble jobs to use a new instance for the success cache, which is empty # [[phab:T383243|T383243]] [[phab:T421114|T421114]]
* 12:17 hashar: Migrated Castor from integration-castor05 to integration-castor06. Updated CASTOR_HOST in Jenkins and moved the Cinder volume to the new instance # [[phab:T421114|T421114]]
* 11:14 hashar: Added Bookworm based Jenkins agents to the pool Hostnames 1090, 1091, 1092 and 1093 # [[phab:T421114|T421114]]
* 10:09 hashar: Added Bookworm based Jenkins agents to the pool Hostnames 1083 to 1089 # [[phab:T421114|T421114]]
* 07:23 hashar: CI Jenkins: removed `blubber` label from all agents after having moved PipelineLib to use the `Docker` label {{!}} [[phab:T422234|T422234]]
=== 2026-04-06 ===
* 16:01 dancy: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1268239
=== 2026-04-03 ===
* 20:17 bd808: Unblock 2.54.0.0/16 ([[phab:T422238|T422238]])
* 17:25 bd808: Unblock 31.18.0.0/16 ([[phab:T422245|T422245]])
* 17:18 bd808: Unblock 2.54.128.0/19 ([[phab:T422238|T422238]])
* 16:18 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1264649 "add Python 3.14 to pywikibot jobs and separate lint tests" {{!}} [[phab:T421723|T421723]]
* 09:26 hashar: integration: nuked pywikibot/core pre-commit cache # [[phab:T422242|T422242]]
* 09:15 hashar: Added Bookworm based Jenkins agents to the pool with label `Docker`. Hostnames are `integration-agent-docker-107*` # [[phab:T421114|T421114]]
* 02:47 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1267398
=== 2026-04-02 ===
* 16:50 thcipriani: restart jenkins
* 15:15 bd808: Unblock 82.216.0.0/16 ([[phab:T421508|T421508]])
* 15:07 bd808: Unblock 95.90.0.0/15 ([[phab:T421485|T421485]])
* 11:19 James_F: Zuul: [oojs/ui] Drop ooui-ruby2.7-rake job, we're abandoning Ruby use there
=== 2026-04-01 ===
* 22:01 bd808: Unblock 109.144.0.0/12 ([[phab:T422019|T422019]])
* 20:16 bd808: Unblock 93.192.0.0/10 ([[phab:T421894|T421894]])
* 19:25 dancy: Updating buildkitd to v0.29.0 in gitlab-cloud-runners (prod) ([[phab:T415284|T415284]])
* 17:57 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/97 ([[phab:T420441|T420441]])
* 17:39 bd808: Unblock 94.134.0.0/15 ([[phab:T421866|T421866]])
* 16:31 dancy: Upgrade buildkit to 0.29.0 in staging gitlab-cloud-runners ([[phab:T415284|T415284]])
* 10:47 taavi: integration-castor05: free up a bit of disk space by deleting cache for AhoCorasick/ CLDRPluralRuleParser/ HtmlFormatter/ RelPath/ RunningStat/ IPSet/
=== 2026-03-30 ===
* 22:01 bd808: Unblock 78.20.0.0/14 ([[phab:T421586|T421586]])
* 21:04 bd808: Unblock 95.88.0.0/15 ([[phab:T421774|T421774]])
* 20:49 bd808: Unblock 95.89.191.0/24 ([[phab:T421774|T421774]])
* 20:29 bd808: Unblock 73.162.0.0/16 ([[phab:T421549|T421549]])
* 13:10 hashar: gerrit: abandon mediawiki/core changes that are 2+years old and are attached to a task (`Bug: Txxxx`)
* 11:37 hashar: Reloaded Zuul to to add 3 persons to the allow list
* 10:43 James_F: Docker: Re-pushing to try to create quibble-coverage 1.16.0-s2
=== 2026-03-27 ===
* 21:00 James_F: Docker: [quibble-bullseye] Drop Python 2 from images
* 11:28 hashar: deployment-prep: removed block for `143.176.0.0/15` and blocked subblock `143.176.0.0/16` instead. This unblocks `143.177.0.0/16` # [[phab:T421420|T421420]]
* 00:18 bd808: Unblock 95.90.238.0/23 ([[phab:T421447|T421447]])
=== 2026-03-26 ===
* 21:25 bd808: Unblock 89.240.0.0/15 ([[phab:T421364|T421364]])
* 21:09 brennen: patchdemo: deploy to production for https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/312
=== 2026-03-25 ===
* 20:41 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256318 [[phab:T421283|T421283]]
* 15:46 dancy: Migrated gitlab-cloud-runners (prod) from nginx-ingress to traefik ([[phab:T420743|T420743]])
* 15:32 dancy: Migrated gitlab-cloud-runners (staging) from nginx-ingress to traefik ([[phab:T420743|T420743]])
* 10:01 hashar: Updating tox Jenkins jobs to add support for Python 3.14 {{!}} https://gerrit.wikimedia.org/r/1260632 {{!}} [[phab:T421209|T421209]]
* 08:40 codders: integration: integration-castor05: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20/
=== 2026-03-24 ===
* 19:40 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1255746
* 15:34 brennen: gitlab1004: manual test run of `configure-projects` with cleared issue allowlist ([[phab:T412882|T412882]])
* 15:26 bd808: Unblock 47.194.0.0/16 ([[phab:T421127|T421127]])
* 12:53 hashar: integration: deleted old Puppet 5 compiler agents from Jenkins ( pcc-worker1014.puppet-diffs.eqiad1.wikimedia.cloud , pcc-worker1015.puppet-diffs.eqiad1.wikimedia.cloud , pcc-worker1016.puppet-diffs.eqiad1.wikimedia.cloud ) # [[phab:T367399|T367399]]
* 07:42 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1259755
=== 2026-03-23 ===
* 15:28 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 90272
=== 2026-03-22 ===
* 14:52 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1258082
* 01:00 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256488
=== 2026-03-21 ===
* 08:10 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256962
* 07:48 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256946
=== 2026-03-20 ===
* 21:21 bd808: Unblock 103.159.218.0/24 ([[phab:T420530|T420530]])
* 14:59 James_F: Zuul: [mediawiki/extensions/AbuseFilter] Add dependency on CodeMirror, for [[phab:T399673|T399673]]
=== 2026-03-19 ===
* 16:54 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1255777
* 16:01 Krinkle: Hoist l10n-bot rights from labs/tools parent to labs parent to reduce duplication in other labs/ repos
* 15:50 Krinkle: Create labs/xtools repo (branch: main, parent: labs, owner: labs-xtools), ref [[phab:T402086|T402086]]
=== 2026-03-18 ===
* 21:11 dcausse: [[phab:T403775|T403775]]: reindexing all wikis to enable new sorting options
* 21:08 dcausse: restarting opensearch on deployment-cirrussearch(12{{!}}13{{!}}14) instances to pickup new plugin versions
* 14:56 James_F: Zuul: Handle wmf/next the same way as wmf/branch_cut_pretest
* 14:52 James_F: Zuul: [GrowthExperiments] drop duplicate VisualEditor dep
* 14:52 James_F: Zuul: [search/*] Add experimental Java 25 jobs
=== 2026-03-17 ===
* 22:50 James_F: Zuul: [mediawiki/extensions/JsonForms] Add quibble jobs
* 21:27 James_F: Zuul: search: Update opensearch plugins for Java 11/17, for [[phab:T420407|T420407]]
* 20:20 bd808: Resize deployment-sessionstore06 from g4.cores1.ram2.disk20 to g4.cores2.ram4.disk20 ([[phab:T415021|T415021]])
* 16:43 James_F: Zuul: [BlueSpicePermissionManager] Add …ConfigManager & …UserManager deps
* 14:36 James_F: Zuul: [mediawiki/extensions/ArticleGuidance]: Add SpamBlacklist as phan dep, for [[phab:T420015|T420015]]
=== 2026-03-13 ===
* 13:59 andrewbogott: deleting ptr record 117.0.16.172.in-addr.arpa. -- accidental duplicate for deployment-kafka-logging01.deployment-prep.eqiad1.wikimedia.cloud
* 13:04 elukey: re-create kafka-logging-01 in deployment-prep on trixie and Kafka 3.7 (was running on buster)
* 09:13 elukey: upgrade kafka-jumbo and kafka-main to Confluent 7.7 in deployment-prep (pre-requisite before being able to upgrade to Trixie)
=== 2026-03-12 ===
* 21:23 bd808: Hard reboot deployment-sessionstore06 ([[phab:T415021|T415021]])
* 01:14 James_F: Docker: [helm-linter] Bump for Envoy 1.35.9, for [[phab:T419637|T419637]]
=== 2026-03-11 ===
* 16:48 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/MetricsPlatform # [[phab:T417568|T417568]]
* 16:47 James_F: Zuul: [mediawiki/extensions/MetricsPlatform] Archive, for [[phab:T416865|T416865]]
* 11:12 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1250529 "inference-services: Split policy violation CI into separate model jobs." - [[phab:T418832|T418832]]
=== 2026-03-10 ===
* 17:39 dduvall: deployed reggie v1.18.0 to gitlab-cloud-runner production
* 17:11 hashar: Updated MediaWiki coverage jobs so that they now keep "Generate a local configuration by running `composer phpunit:config`" message # [[phab:T419073|T419073]]
* 16:41 dduvall: deployed reggie v1.18.0 to gitlab-cloud-runner staging
* 08:21 codders: integration: integration-castor05: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20
=== 2026-03-09 ===
* 21:53 bd808: Reboot deployment-shellbox01 on the off chance that is makes the new permissions error go away ([[phab:T419440|T419440]])
* 13:13 James_F: Zuul: [mediawiki/extensions/WikiShare] Mark as archived, for [[phab:T413589|T413589]]
* 13:11 James_F: Zuul: [mediawiki/extensions/Memento] Mark as archived, for [[phab:T369991|T369991]]
* 13:10 James_F: Zuul: [mediawiki/extensions/QuickGV] Mark as archived, for [[phab:T413348|T413348]]
* 13:10 James_F: Zuul: [mediawiki/extensions/SemanticImageInput] Mark as archived, for [[phab:T413588|T413588]]
* 13:09 James_F: Zuul: [mediawiki/extensions/SidebarDonateBox] Mark as archived, for [[phab:T413587|T413587]]
* 13:07 James_F: Zuul: [mediawiki/extensions/SemanticSifter] Mark as archived, for [[phab:T413586|T413586]]
* 13:06 James_F: Zuul: [mediawiki/extensions/GoogleAdSense] Mark as archived, for [[phab:T413585|T413585]]
* 13:04 James_F: Zuul: [mediawiki/extensions/SecurityAPI] Mark as archived, for [[phab:T418008|T418008]]
* 12:50 James_F: Zuul: [mediawiki/extensions/CheckUser] Add DiscussionTools dependency
* 12:50 James_F: Zuul: [mediawiki/skins/MinervaNeue] Add dependencies for TestKitchen
* 10:40 hashar: gerrit: mediawiki/vendor: converted `es6` and `es710` branches to tags # [[phab:T417804|T417804]]
* 09:24 hashar: Updating Quibble jobs to 1.16.0 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1248880 {{!}} [[phab:T417399|T417399]] [[phab:T417409|T417409]] [[phab:T418461|T418461]]
* 09:15 hashar: updating all CI Jenkins jobs using `./jjb-update`
=== 2026-03-06 ===
* 19:46 James_F: Zuul: [mediawiki/services/geoshapes] Mark as archived, for [[phab:T418372|T418372]]
* 16:37 hashar: Building Docker images for Quibble 1.16.0
* 16:31 hashar: Tag Quibble 1.16.0 @ {{Gerrit|0b9db5fe3cabb2cec0b5d44e128bafa917b3b895}} # [[phab:T417399|T417399]] [[phab:T417409|T417409]] [[phab:T418461|T418461]]
* 12:32 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1248411 "jjb, Zuul: vary Wikibase Selenium for release branches" {{!}} [[phab:T418797|T418797]]
* 12:12 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1248409/ "jjb, Zuul: rename wikibase-selenium job for clarity" {{!}} [[phab:T418797|T418797]]
=== 2026-03-05 ===
* 14:41 James_F: Zuul: [mediawiki/skins/MinervaNeue] Add TestKitchen as a dependency for [[phab:T418053|T418053]]
* 08:01 hashar: Reloaded Zuul to rename wikibase-client / wikibase-repo jobs {{!}} https://gerrit.wikimedia.org/r/1238317
* 00:04 James_F: Docker: [quibble-coverage] Use local PHPUnit config, for [[phab:T345481|T345481]]
=== 2026-03-04 ===
* 21:16 James_F: Zuul: [mediawiki/core] Make PHP 8.5 voting on master branch, for [[phab:T411814|T411814]]
* 21:10 James_F: Zuul: [mediawiki/vendor] Make PHP 8.5 voting on master branch, for [[phab:T411814|T411814]]
* 19:48 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/96 ([[phab:T419004|T419004]])
* 18:50 James_F: Revert "Zuul: [mediawiki/extensions/MobileFrontend] Add ParserMigration dependency", for [[phab:T419043|T419043]]
* 16:23 James_F: Zuul: [mediawiki/services/parsoid] Make PHP 8.4 voting
* 15:37 James_F: Docker: [rake-ruby2.7] Add libffi-dev too, for [[phab:T418463|T418463]]
* 13:59 James_F: Docker: [rake-ruby2.7] Add ruby-ffi for [[phab:T418463|T418463]]
* 13:54 hashar: SIGKILL Zuul cause it can't gracefully stop most probably due to being locked attempting to report back to Gerrit # [[phab:T419009|T419009]]
* 13:49 hashar: Stopping Zuul # [[phab:T419009|T419009]]
* 13:41 hashar: Took a Zuul stack dump on contint1002.wikimedia.org using SIGUSR1 # [[phab:T419009|T419009]]
=== 2026-03-03 ===
* 23:52 James_F: Zuul: [mediawiki/extensions/WikimediaMessages] Drop MetricsPlatform phan dep
* 23:52 James_F: Zuul: [mediawiki/extensions/WikimediaEvents] Drop MetricsPlatform phan dep
=== 2026-03-02 ===
* 22:13 James_F: Zuul: Enforce PHP 8.4 in MW extensions and skins for development branch, for [[phab:T386108|T386108]]
* 14:05 James_F: Zuul: [mediawiki/extensions/MobileFrontend] Add ParserMigration dependency, for [[phab:T415451|T415451]]
* 13:48 James_F: Zuul: […/WikimediaEvents] Drop LoginNotify dependency, now unused, for [[phab:T404334|T404334]]
* 10:16 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/15.8.2/ # [[phab:T418718|T418718]]
=== 2026-02-28 ===
* 21:33 hashar: gerrit: triggering replication to GitHub for all of `mediawiki/skins` # [[phab:T418675|T418675]]
* 21:33 hashar: gerrit: triggering replication to GitHub for all of `mediawiki/extensions` # [[phab:T418675|T418675]]
=== 2026-02-27 ===
* 15:53 dancy: Updating gitlab-cloud-runners (staging and prod) to gitlab-runner 18.9.0.
=== 2026-02-26 ===
* 20:16 James_F: Zuul: Provide a custom, high-priority pipeline just for puppet compiler [[phab:T414621|T414621]]
* 19:32 James_F: Docker: Bump all the PHPs.
* 13:40 hashar: Deployed Jenkins job https://integration.wikimedia.org/ci/job/wikibase-selenium/ # [[phab:T287582|T287582]]
* 00:13 dduvall: forcing replacement of buildkitd helm release in gitlab-cloud-runner prod cluster due to dependency on removed k8s secret ([[phab:T416260|T416260]])
=== 2026-02-25 ===
* 23:50 dduvall: deploying https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/552 to gitlab-cloud-runner production cluster ([[phab:T416260|T416260]])
* 14:07 James_F: Zuul: [mediawiki/extensions/CommunityRequests] Add TemplateData dependency, for [[phab:T401638|T401638]]
* 00:08 jeena: no-op testing updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/95
=== 2026-02-24 ===
* 15:55 brennen: devtools: test deploy phab/phorge to test instance ([[phab:T418256|T418256]])
=== 2026-02-23 ===
* 23:07 jeena: Updated development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/92
* 22:43 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/92
* 22:12 bd808: Unblock 191.80.192.0/18 ([[phab:T418132|T418132]])
* 20:26 hashar: Deleted "replication-upstream" Grafana dashboard in favor of a copy/new "replication" one. https://grafana.wikimedia.org/d/RFLS1GsWk/replication-upstream , replaced it by https://grafana.wikimedia.org/d/d4a4da73-c27f-4ce6-a9e5-ab84dd7a4ebb/replication
* 16:29 James_F: Zuul: [3d2png] Add basic Node CI at version 20
=== 2026-02-20 ===
* 21:47 bd808: Unblock 168.184.84.0/24 ([[phab:T418020|T418020]])
* 17:13 bd808: Unblock 122.187.64.0/18 ([[phab:T417964|T417964]])
* 14:35 James_F: Zuul: [mediawiki/extensions/Monstranto] Move out of Wikimedia prod section
=== 2026-02-19 ===
* 18:34 bd808: Unblock 181.98.0.0/16 ([[phab:T417890|T417890]])
* 17:21 James_F: Zuul: [mediawiki/extensions/WikimediaEvents] Add AbuseFilter as a dependency, for [[phab:T417799|T417799]]
* 13:22 hashar: Reloaded Zuul to archive the Cergen repository {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1240688 {{!}} [[phab:T417887|T417887]]
=== 2026-02-18 ===
* 20:17 jeena: Updating development images on contint primary for [[phab:T415922|T415922]]
* 19:44 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240360
* 18:40 bd808: Unblock 46.59.0.0/17 ([[phab:T417747|T417747]])
* 17:05 hashar: Regenerating Jenkins jobs with JJB based on https://gerrit.wikimedia.org/r/c/integration/config/+/1240254/
* 17:04 hashar: Added EXT_DEPENDENCIES to Quibble Jenkins jobs parameters so we can manually trigger them from the Web UI using a different set of deps # https://gerrit.wikimedia.org/r/c/integration/config/+/1240254/
* 16:30 hashar: Triggered https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/ with empty Zuul parameters introduced by https://gerrit.wikimedia.org/r/1240333 {{!}} https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/4893/console
* 15:43 James_F: Zuul: [mediawiki/extensions/ReadingLists] Add EventBus dependency for [[phab:T417706|T417706]]
* 12:15 hashar: zuul-1001.zuul3.eqiad1.wikimedia.cloud: added keepalive=20 to the scheduler Gerrit driver and restarted scheduler container # [[phab:T417497|T417497]]
* 06:58 jeena: Updating development images on contint primary for [[phab:T415922|T415922]]
=== 2026-02-17 ===
* 23:37 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240081
* 23:20 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240078
* 15:58 brennen: deployed latest phab/phorge wmf/stable to devtools test instance ([[phab:T417657|T417657]])
* 09:01 hashar: Reloaded Zuul to enable php 8.5 testing on utfnormal, php-session-serializer, wikipeg, mediawiki/libs/Dodo, mediawiki/libs/UUID, testing-access-wrapper and translatewiki # [[phab:T406326|T406326]]
=== 2026-02-16 ===
* 15:27 hashar: Manually cleaned some old workspaces on integration-agent-docker-1042
=== 2026-02-12 ===
* 20:07 James_F: Zuul: Enable PHP 8.5 jobs for most MW libraries, for [[phab:T406326|T406326]]
* 19:33 James_F: Docker: [php83] Re-build with upstream's new 8.3.30 release and cascade
* 19:31 James_F: Zuul: Add PHP 8.5 CI job to various things noted as blocked by Phan, for [[phab:T410941|T410941]], [[phab:T406326|T406326]]
* 16:35 Krinkle: Disable publishing noise on tasks from repos Bcp47, clover-diff, ScopedCallback, and IDLeDOM. Ref [[phab:T143162|T143162]]
* 15:53 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/87
* 11:21 James_F: Zuul: [mediawiki/libs/shellbox] Add direct Phan job, for [[phab:T416064|T416064]]
=== 2026-02-10 ===
* 20:16 dancy: Rebooted k3s.catalyst-dev (it was unresponsive, but the reboot hasn't helped)
=== 2026-02-09 ===
* 21:58 James_F: Zuul: [mediawiki/tools/phan] Add PHP 8.5 CI job, for [[phab:T410941|T410941]]
* 19:46 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1238006 [[phab:T415680|T415680]]
* 11:51 James_F: Zuul: [mediawiki/extensions/ReadingLists] Drop MetricsPlatform dependency, for [[phab:T414435|T414435]]
=== 2026-02-05 ===
* 17:58 James_F: Zuul: […/WikimediaCustomizations] Add six new dependencies for [[phab:T404334|T404334]]
* 15:35 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1237254
* 15:18 James_F: Zuul: […/OATHAuth] Add dependency and phan dependency on CentralAuth
=== 2026-02-04 ===
* 12:54 James_F: Zuul: [mediawiki/extensions/Petition] Add CLDR dependency
* 10:03 hashar: Restarted Jenkins on releases2003.codfw.wmnet
=== 2026-02-02 ===
* 21:17 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1234926 "re-enable master jobs for some BlueSpice repos - [[phab:T403196|T403196]]"
* 21:05 bd808: Unblock 85.146.0.0/17 ([[phab:T416079|T416079]])
* 19:47 James_F: Zuul: […/WikimediaCustomizations] Add cldr phan dependency, for [[phab:T404334|T404334]]
* 17:33 bd808: Unblock 188.188.0.0/15 ([[phab:T416095|T416095]])
* 17:26 bd808: Unblock 85.94.84.0/22 ([[phab:T416105|T416105]])
* 17:09 bd808: Unblock 94.234.0.0/16 ([[phab:T416165|T416165]])
* 16:51 dancy: Update gitlab-runners to alpine-v18.6.6 ([[phab:T415214|T415214]])
* 16:27 bd808: Unblock 47.231.208.0/21 ([[phab:T416010|T416010]])
* 11:39 James_F: Zuul: […/WikimediaCustomizations] Add five new phan dependencies, for [[phab:T404334|T404334]]
* 09:45 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 58532, 58557
=== 2026-01-31 ===
* 21:49 James_F: Deleted Jenkins's job entry for castor-save-workspace-cache {{Gerrit|6193776}} and this seems to have unstuck things for [[phab:T416078|T416078]]?
* 21:45 James_F: Running `sudo systemctl restart jenkins` on contint for [[phab:T416078|T416078]]
* 21:44 James_F: Fighting [[phab:T416078|T416078]], took integration-castor-5 offline, disconnected, sshed in to kill threads, then reconnected; no change in aspect.
* 19:03 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1235380
=== 2026-01-28 ===
* 21:26 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/WebAuthn # [[phab:T415832|T415832]]
* 21:11 bd808: Unblock 181.160.0.0/15 & 186.40.128.0/17 ([[phab:T415820|T415820]])
* 17:01 bd808: Unblock 102.182.0.0/16 ([[phab:T415782|T415782]])
=== 2026-01-27 ===
* 16:45 James_F: Zuul: Switch skin-quibble template with identical extension-quibble, for [[phab:T402398|T402398]]
* 16:18 James_F: Zuul: [ArticleGuidance] mention it will be in production
* 15:55 James_F: Docker: [quibble-bullseye] Update to Quibble 1.15.0
* 15:12 James_F: Docker: [quibble-coverage] Pass PHPUnit config location explicitly, for [[phab:T395470|T395470]]
* 09:18 hashar: integration: on integration-castor05, deleted caches for old MediaWiki branches
* 09:15 hashar: integration: on pkgbuilder instances, removed Buster cow images, aptcache and hooks. `sudo cumin --force -p 0 'name:pkgbuilder' 'rm -fR /srv/pbuilder/<nowiki>{</nowiki>base-buster-amd64.cow,hooks/buster,aptcache/buster-amd64<nowiki>}</nowiki>'` # [[phab:T397209|T397209]]
* 09:14 hashar: integration: cleaned up old workspaces under /srv/jenkins/workspace
=== 2026-01-26 ===
* 23:27 bd808: Unblock 66.130.0.0/15 ([[phab:T415596|T415596]])
* 22:52 bd808: Unblock 45.16.0.0/12 ([[phab:T415467|T415467]])
* 14:46 hashar: gerrit: changed `operations/software/permissions` project type from `CODE` to `PERMISSIONS` by pointing `HEAD` to `refs/meta/config`
=== 2026-01-22 ===
* 17:36 James_F: Docker: [quibble-coverage] Stop using legacy PHPUnit entrypoint ([[phab:T395470|T395470]]) & Stop excluding Dump/ParserFuzz/Stub groups ([[phab:T415230|T415230]])
* 15:11 James_F: Zuul: [mediawiki/extensions/Math] Add a standalone job, for [[phab:T415230|T415230]]
=== 2026-01-20 ===
* 20:38 bd808: Cherry picked https://gerrit.wikimedia.org/r/c/operations/puppet/+/1229186 ([[phab:T415113|T415113]])
* 19:05 bd808: Rebooted deployment-cache-text08 to see if the mystery haproxy startup failure would go away ([[phab:T415100|T415100]])
* 18:50 bd808: Unblock 152.7.0.0/16 ([[phab:T415100|T415100]])
=== 2026-01-17 ===
* 23:32 ori: beta-scap with `php_l10n: true` completed successfully: https://integration.wikimedia.org/ci/view/Beta/job/beta-scap-sync-world/241466/console. PHP l10n files generated. Reverted local change to scap.cfg.
* 23:26 ori: Temporarily set `php_l10n: true` on deployment-deploy04:/etc/scap.cfg to see if next scap succeeds.
=== 2026-01-16 ===
* 16:33 dancy: Deleting deployment-mx03.deployment-prep ([[phab:T412975|T412975]])
=== 2026-01-15 ===
* 14:50 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/ArticleSummaries/ # [[phab:T413232|T413232]]
=== 2026-01-14 ===
* 17:14 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1226907
* 16:27 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1226893
* 15:57 bd808: Unblock 190.60.63.0/24 ([[phab:T414541|T414541]])
=== 2026-01-13 ===
* 15:04 James_F: Zuul: Make quibble-for-mediawiki-core-vendor-mysql-php84 voting, for [[phab:T386108|T386108]]
=== 2026-01-12 ===
* 21:33 zabe: zabe@deployment-mwmaint03:~$ foreachwiki migrateLinksTable.php --table imagelinks # [[phab:T413668|T413668]]
* 21:06 bd808: Unblock 66.81.168.0/21 ([[phab:T414303|T414303]])
* 17:42 dancy: Turned off instance deployment-prep.deployment-mx03
* 11:44 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 46331, 46344
=== 2026-01-10 ===
* 21:48 taavi: reload zuul for https://gerrit.wikimedia.org/r/1224782
* 00:25 bd808: Unblock 91.160.0.0/12 ([[phab:T414190|T414190]])
=== 2026-01-09 ===
* 17:33 thcipriani: re-enabling beta update jobs after test bad extension-list [[phab:T411516|T411516]]
* 17:09 thcipriani: disabling beta update jobs to test bad extension-list [[phab:T411516|T411516]])
=== 2026-01-08 ===
* 21:30 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1224815 [[phab:T414136|T414136]]
* 18:24 bd808: Unblock 89.80.0.0/12 ([[phab:T414113|T414113]])
* 15:55 dancy: Upgrading gitlab-runner to v18.5.0 on gitlab-cloud-runners. ([[phab:T414053|T414053]])
=== 2026-01-07 ===
* 23:17 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1082574 https://gerrit.wikimedia.org/r/1224157 https://gerrit.wikimedia.org/r/1224159
* 23:12 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/896311 [[phab:T27482|T27482]]
* 23:06 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1224218
* 17:34 James_F: Zuul: Add new extensions: IssueTrackerLinks, PreviewLinks, and WikiRAG
* 17:34 James_F: Zuul: [labs/tools/heritage] Point to the task to drop 8.1 testing
* 15:09 James_F: Zuul: [labs/tools/heritage] Add testing in PHP 8.2+, not just PHP 8.1
* 15:03 James_F: Zuul: Even for extension-broken, don't offer PHP 8.1 testing
* 15:02 James_F: Zuul: Move quibble experimental sqlite/postgres tests to PHP 8.3
=== 2026-01-06 ===
* 16:57 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1223690 [[phab:T411814|T411814]]
* 16:16 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1223189 [[phab:T411814|T411814]]
* 00:30 bd808: Unblock 85.134.128.0/17 ([[phab:T413755|T413755]])
* 00:02 bd808: Unblock 89.166.128.0/17 ([[phab:T413702|T413702]])
=== 2026-01-05 ===
* 23:57 bd808: Unblock 185.233.104.0/22 ([[phab:T413472|T413472]])
* 23:51 bd808: Unblock 45.62.112.0/21 ([[phab:T413079|T413079]])
* 23:44 bd808: Unblock 85.134.200.0/21 ([[phab:T413067|T413067]])
* 19:03 dancy: Updated buildkitd to v0.26.3 in gitlab-cloud-runners
* 14:27 taavi: reload zuul for {{Gerrit|1223191}}
* 13:57 James_F: Zuul: [mediawiki/php/wmerrors] Enable PHP 8.5 testing, for [[phab:T410921|T410921]]
=== 2026-01-03 ===
* 17:59 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1222709 https://gerrit.wikimedia.org/r/1220388 https://gerrit.wikimedia.org/r/1219140
=== 2026-01-02 ===
* 17:10 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1222597
=== 2026-01-01 ===
* 02:34 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1221644
<noinclude>'''Server Admin Log''' logged from {{IRC|wikimedia-releng}} for [[Nova Resource:Deployment-prep|Beta Cluster]], [[mw:Continuous integration|Continuous integration]] and various other Release Engineering projects.</noinclude>
{{SAL-archives/Release Engineering}}
<noinclude>[[Category:SAL]]</noinclude>
sldem9ffpqq09zzg25bdbr5k13y3jub
2461135
2461128
2026-09-26T23:33:58Z
Stashbot
7414
Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1345303
2461135
wikitext
text/x-wiki
=== 2026-09-26 ===
* 23:33 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1345303
* 20:11 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1345288
=== 2026-09-25 ===
* 15:43 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 15:23 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 15:20 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.1-do.6 to 1.35.7-do.5 ([[phab:T439035|T439035]])
* 15:05 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.1-do.6 to 1.35.7-do.5 ([[phab:T439035|T439035]])
* 14:05 hashar: Tag Quibble 1.22.0 @ {{Gerrit|0d0bce00eedf44dd235703cb3eb1f60083d64752}} # [[phab:T437752|T437752]] [[phab:T426684|T426684]]
=== 2026-09-24 ===
* 17:02 dancy: Upgraded gitlab-cloud-runners (staging) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 16:01 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.7-do.5 to 1.36.3-do.5 ([[phab:T439035|T439035]])
* 15:37 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.1-do.6 to 1.35.7-do.5 ([[phab:T439035|T439035]])
* 15:30 dancy: Update gitlab-runner image to v19.2.6 ([[phab:T439037|T439037]])
=== 2026-09-23 ===
* 22:06 Southparkfan: decommission Beta Cluster PHP 8.3 hosts deployment-mediawiki13, deployment-mediawiki14, deployment-jobrunner05, deployment-mwmaint03 - [[phab:T435393|T435393]]
* 21:46 Southparkfan: cherry-pick https://gerrit.wikimedia.org/r/c/operations/puppet/+/1344336 to Puppetserver - [[phab:T435393|T435393]]
* 20:37 hashar: gerrit: deleted https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1344362 duplicate Change-Id {{Gerrit|I67dd7b3c879049a1c98586c365c297eac150b963}} of https://gerrit.wikimedia.org/r/c/mediawiki/core/+/1329697
* 20:26 dancy: Upgrading istio to 1.30.5 on gitlab-cloud-runners (staging) ([[phab:T439035|T439035]])
* 20:13 hashar: gerrit: reindexing has been completed
* 19:38 hashar: gerrit: started reindexing for accounts, group and projects indices
* 19:36 hashar: gerrit: force online reindexing of the change index over ssh hitting gerrit1003 (switch over). `gerrit index start changes --force`
* 12:58 awight: add seanleong-wmde to deployment-prep
* 12:56 awight: manually run puppet agent
=== 2026-09-22 ===
* 22:43 Southparkfan: add wmgRedisLockPassword to PrivateSettings.php - [[phab:T436480|T436480]]
* 22:40 Southparkfan: create empty /etc/helmfile-defaults/mediawiki/release on -deploy04 and -deploy06 to unbreak Scap sync-masters - [[phab:T435393|T435393]]
* 21:22 Southparkfan: decommission deployment-cumin-3 - [[phab:T436470|T436470]]
* 21:20 Southparkfan: add deployment-deploy06 to scap::dsh::scap_masters and deployment_hosts - [[phab:T435393|T435393]]
* 20:53 Southparkfan: remove puppetserver cherry-pick for [[phab:T428052|T428052]], running newer version of HAProxy nowadays
* 20:08 Southparkfan: add deployment-jobrunner06 to 'jobrunner' dsh group - [[phab:T438656|T438656]]
* 18:17 dancy: Remove deployment-mwmaint03.deployment-prep.eqiad1.wikimedia.cloud from scap::dsh::groups.mediawiki-installation.host in deployment-prep project puppet config ([[phab:T435393|T435393]]) to stop wmf-beta-update-all spam.
* 08:53 hashar: Restarted CI Jenkins on contint1003 to apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1343675 "jenkins: exit the JVM on OutOfMemoryError" # [[phab:T435791|T435791]]
* 07:21 awight: Purge opcache on deployment-jobrunner06.deployment-prep
=== 2026-09-21 ===
* 15:54 hashar: gerrit: on mediawiki/extensions/QuickSurveys deleted branch `master-backup` which was pointing at {{Gerrit|1ad54717c059bcd49093d902eab2c098b4efe91a}} (which is in `master`)
* 13:56 awight: Purging opcache on deployment-jobrunner06
* 11:37 hashar: Build docker-registry.wikimedia.org/releng/ajv:0.4.1-s5 (upgrade Node from 24.18.0 to 26.8.2) Used by PipelineLib # [[phab:T438085|T438085]]
=== 2026-09-18 ===
* 15:27 James_F: CI Node reverted to 24.
* 14:46 James_F: Zuul: [labs/tools/wdaudiolex-be] Install tox CI
* 14:01 James_F: Zuul: Remaining Node 24 -> 26 migrations
* 13:48 James_F: Zuul: Migrate MediaWiki-land independent CI Node from 24 to 26, for [[phab:T438085|T438085]]
=== 2026-09-17 ===
* 21:27 hashar: Re running `postmerge` for https://gerrit.wikimedia.org/r/c/mediawiki/services/wikifeeds/+/1342700 for [[phab:T438379|T438379]] # `zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/services/wikifeeds --change {{Gerrit|1342700}},1`
* 16:24 brett: delete deployment-cache, deployment-cache-text, and deployment-cache-upload prefixes - [[phab:T436468|T436468]]
* 15:59 hashar: gerrit: manually deleted old wmf branches from VisualEditor/VisualEditor # [[phab:T438373|T438373]]
* 15:09 hashar: gerrit: convert old REL branches on VisualEditor/VisualEditor ( [[phab:T380841|T380841]] [[phab:T428864|T428864]] [[phab:T428911|T428911]] ) using: for branch in REL1_25 REL1_26 REL1_27 REL1_28 REL1_29 REL1_30 REL1_31 REL1_32 REL1_33 REL1_34 REL1_35 REL1_36 REL1_37 REL1_38 REL1_39 REL1_40 REL1_41 REL1_42 REL1_44; do ./.tox/make-release/bin/python ./make-release/branch.py --delete --abandon --bundle ve $branch; done;
* 15:00 hashar: gerrit: converting REL1_44 branches to tags to formally EOL REL1_44 {{!}} [[phab:T428911|T428911]]
* 13:13 hashar: Updated plugins on the CI Jenkins
* 07:29 codders: rebuilt and restarted phpunit-results-cache server
=== 2026-09-16 ===
* 06:15 hashar: Updating node jobs from NodeJS 26.4.0 to 27.8.2 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1342056
=== 2026-09-15 ===
* 20:46 James_F: Docker: [quibble-bookworm, node24-test, node26-test] Drop jsduck etc. for [[phab:T363905|T363905]]
* 20:39 James_F: Zuul: […/OOJsUIAjaxLogin] Drop the JavaScript documentation job, for [[phab:T391706|T391706]]
* 19:09 Southparkfan: add deployment-cp-text09 and deployment-cp-upload09 to 'cache_hosts' hiera key - [[phab:T436468|T436468]]
* 18:33 James_F: Zuul: [labs/tools/ldap] Archive repository, for [[phab:T438076|T438076]]
* 18:32 James_F: Zuul: [integration/gear] Archive repository, for [[phab:T289512|T289512]] and [[phab:T438076|T438076]]
* 18:29 James_F: Zuul: [research/landing-page] Archive repository, for [[phab:T438076|T438076]]
* 18:28 James_F: Zuul: [cloud/toolforge/*] Archive two Toolforge repositories, for [[phab:T438076|T438076]]
* 18:26 James_F: Zuul: [node-rdkafka-factory, node-rdkafka-statsd] Archive repos, for [[phab:T366611|T366611]] and [[phab:T438076|T438076]].
* 18:24 James_F: Zuul: [operations/container/miscweb] Archive repository, for [[phab:T438076|T438076]]
* 18:22 James_F: Zuul: [mediawiki/services/recommendation-api] Archive service, for [[phab:T429123|T429123]] and [[phab:T438076|T438076]]
* 18:19 James_F: Zuul: [wikidata/query-builder, wikidata/query/gui] Archive repos, for [[phab:T438076|T438076]]
* 18:16 James_F: Zuul: [mediawiki/services/wikispeech/*] Archive the services, for [[phab:T344741|T344741]] and [[phab:T438076|T438076]]
* 10:38 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # corrupt Cypress cache? [[phab:T438002|T438002]]
* 10:00 hashar: Updating node based Jenkins jobs for https://gerrit.wikimedia.org/r/c/integration/config/+/1341713 {{!}} update npm jobs to drop debug logs from cache # [[phab:T437376|T437376]] [[phab:T426741|T426741]]
* 05:19 phedenskog: devel-stats upgraded datasette-dashboards from 0.7.1 to 0.8.0 for releng-data.wmcloud.org
=== 2026-09-14 ===
* 19:51 James_F: Zuul: [mediawiki/extensions/MobileApp] Add VisualEditor phan dep, for [[phab:T437736|T437736]]
* 19:49 Southparkfan: switch Beta Cluster to PHP 8.5 - [[phab:T435393|T435393]]
* 17:18 taavi: relaoding zuul to deploy https://gerrit.wikimedia.org/r/1337606
=== 2026-09-12 ===
* 18:23 Southparkfan: project-wide puppet hiera: remove puppetmaster::geoip::<nowiki>{</nowiki>fetch_private,use_proxy<nowiki>}</nowiki>, puppetdb_host, profile::puppetmaster::common::<nowiki>{</nowiki>command_broadcast,puppetdb_host,puppetdb_hosts<nowiki>}</nowiki>
* 18:14 Southparkfan: project-wide puppet hiera: remove role::puppetmaster::puppetdb::shared_buffers, obsoleted by profile::puppetdb::database::shared_buffers
* 18:09 Southparkfan: project-wide puppet hiera: remove profile::puppetdb::master pointing to former deployment-puppetdb02 (overridden in deployment-puppetdb prefix), profile::puppetdb::slaves (by default already empty) + profile::puppetdb::extra_authorized_hosts (variable does not exist, couldn't find it in historic commits either)
=== 2026-09-11 ===
* 16:29 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1339162
* 06:05 phedenskog: Updating Jenkins jobs from https://gerrit.wikimedia.org/r/c/integration/config/+/1338086 (quibble*, api-testing-*, mwext-phpunit-coverage*): run npm cache verify at most once a day per cache [[phab:T437377|T437377]]
=== 2026-09-10 ===
* 22:01 Southparkfan: decommission deployment-mwlog02 - [[phab:T436469|T436469]]
* 21:08 Southparkfan: masked and stopped stray udp2log.service on deployment-mwlog03, claimed 8420/udp which should have been assigned to udp_tee - [[phab:T436469|T436469]]
* 18:56 James_F: Zuul: [mediawiki/extensions/PersonalDashboard] Add ORES & Wikibase phan deps for [[phab:T436570|T436570]] and [[phab:T437491|T437491]]
* 18:55 James_F: Zuul: [mediawiki/extensions/ArticleGuidance] Add CommunityConfiguration for phan for [[phab:T437586|T437586]]
* 05:49 phedenskog: Updated quibble-for-mediawiki-core postgres and sqlite jobs to run only the phpunit-database stage [[phab:T437134|T437134]]
* 05:20 phedenskog: Updated 134 Quibble Jenkins jobs to drop the duplicated --reporting-url [[phab:T323750|T323750]]
=== 2026-09-09 ===
* 22:36 Southparkfan: point profile::rsyslog::udp_tee::destinations to both deployment-mwlog02 and deployment-mwlog03 8421/udp - [[phab:T436469|T436469]]
* 22:32 Southparkfan: set role::logging::mediawiki::udp2log::monitor: false to unbreak [[phab:T436469|T436469]]
* 16:53 Southparkfan: cpjobqueue: switch back to deployment-jobrunner05 (PHP 8.3) - [[phab:T435393|T435393]]
* 16:32 Southparkfan: adjust cpjobqueue config to temporarily send jobs to deployment-jobrunner06 (PHP 8.5) - [[phab:T435393|T435393]]
* 16:12 Southparkfan: decommission deployment-docker-mathoid02 - [[phab:T436465|T436465]]
* 16:09 thcipriani: deployed some anti-scraper mitigations on beta
* 15:40 komla@cloudcumin1001: END (PASS) - Cookbook wmcs.openstack.quota_increase (exit_code=0) by 30 server-groups ([[phab:T437117|T437117]])
* 15:40 komla@cloudcumin1001: START - Cookbook wmcs.openstack.quota_increase by 30 server-groups ([[phab:T437117|T437117]])
* 11:58 fnegri@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.add_user_to_project (exit_code=0) for user 'denisse' in role 'member'
* 11:58 fnegri@cloudcumin1001: START - Cookbook wmcs.vps.add_user_to_project for user 'denisse' in role 'member'
* 09:42 phedenskog: jjb: update quibble-with-gated-extensions-selenium-php83 to run npm install ahead of the browser tests [[phab:T436376|T436376]]
=== 2026-09-08 ===
* 19:36 Southparkfan: truncate Apache2 forensic log on deployment-mediawiki14, disk almost full
* 19:35 Southparkfan: set profile::mediawiki::httpd::enable_forensic_log to false, to avoid disk exhaustion during high-traffic load
* 16:21 thcipriani: hard reboot deployment-prep:deployment-cache-text08 horizon logs showing OOMs
* 15:42 hashar: Updating Quibble jobs to 1.21.0 # [[phab:T429715|T429715]] [[phab:T303270|T303270]] [[phab:T437134|T437134]] [[phab:T300727|T300727]]
* 15:06 hashar: Tag Quibble @ {{Gerrit|b692707e5213abf5e16894b1fb21cea41cd72c9e}} # [[phab:T429715|T429715]] [[phab:T303270|T303270]] [[phab:T437134|T437134]] [[phab:T300727|T300727]]
=== 2026-09-07 ===
* 19:31 Southparkfan: 19:31 UTC: switch back from PHP 8.5 to PHP 8.3 hosts - [[phab:T435393|T435393]]
* 19:06 Southparkfan: 18:05 UTC: pool mediawiki15 and mediawiki16 (PHP 8.5) as replacements for 13 and 14 (PHP 8.3), smoke test - [[phab:T435393|T435393]]
* 19:03 Southparkfan: banhammer lots of ranges to get Beta Cluster back online; not sure it was very effective, but we seem to be out of the woods
* 17:55 Southparkfan: no disk space left on deployment-mediawiki14, cleared logs in /var/log/apache2/forensic to unbreak
=== 2026-09-04 ===
* 15:52 hashar: integration: remove from Jenkins global config: NPM_CONFIG_AUDIT=false and NPM_CONFIG_FUND=false # [[phab:T437008|T437008]]
* 14:43 hashar: integration: set in Jenkins global config: NPM_CONFIG_AUDIT=false and NPM_CONFIG_FUND=false # [[phab:T437008|T437008]]
=== 2026-09-03 ===
* 23:29 Southparkfan: attach Cinder volume for /srv on deploy06, jobrunner06, mediawiki15/mediawiki16 - [[phab:T435393|T435393]]
* 11:26 hashar: Deleted coverage report for SimilarEditors ( /srv/doc/cover-extensions/SimilarEditors ), extension is being archived # [[phab:T436880|T436880]]
* 08:13 James_F: Zuul: [mediawiki/services/similar-users] Archive service, for [[phab:T368269|T368269]]
* 07:25 hashar: integration: granted sudo access to Phedenskog
* 06:27 hashar: Upgrading CI Jenkins on contint1003 # [[phab:T436812|T436812]]
* 05:15 hashar: Updating tox jobs to change default python from 3.9 to 3.11 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1334023 {{!}} [[phab:T436857|T436857]]
=== 2026-09-02 ===
* 22:21 Southparkfan: decommission deployment-webperf21 - [[phab:T436464|T436464]]
* 17:09 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1333891
* 15:37 hashar: jjb: update Quibble jobs to 1.20.0 # [[phab:T321617|T321617]] [[phab:T432966|T432966]] [[phab:T428642|T428642]] [[phab:T435975|T435975]]
* 15:14 hashar: Building Quibble 1.20.0 images
* 14:52 hashar: Tag Quibble 1.20.0 @ {{Gerrit|c6f13962c27685773fc1fb39d74189b202e0961d}} # [[phab:T321617|T321617]] [[phab:T432966|T432966]] [[phab:T428642|T428642]] [[phab:T435975|T435975]]
=== 2026-09-01 ===
* 21:45 Southparkfan: hiera: switch ATS routing for performance.beta.wmcloud.org from webperf21 to webperf31 - [[phab:T436464|T436464]]
* 20:55 Southparkfan: decommission deployment-webperf22 - [[phab:T436464|T436464]]
* 20:52 Southparkfan: https://gerrit.wikimedia.org/r/c/operations/puppet/+/1333288 cherry-picked on project puppetserver - [[phab:T436464|T436464]]
* 19:59 Southparkfan: decommission deployment-webperf32 - [[phab:T436464|T436464]]
* 17:54 Southparkfan: switch webproxy for wikifeeds-beta to deployment-docker-wikifeeds01 - [[phab:T436462|T436462]]
* 17:47 Southparkfan: switch profile::restbase::citoid_uri and profile::restbase::cxserver_uri to resp. citoid03 and cxserver03, old VMs no longer exist - [[phab:T436619|T436619]]
* 17:43 Southparkfan: revoked Puppet certs for deployment-docker-cxserver02 and deployment-docker-citoid02 - [[phab:T436619|T436619]]
* 14:36 hashar: integration: deleted Cypress from codehealth job after the job learn to instruct Cypress & Puppeter to no more download binary blobs. `rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-codehealth-master-non-voting/Cypress` # [[phab:T427471|T427471]]
=== 2026-08-31 ===
* 22:17 andrewbogott: (log again, mentioned wrong task last time) add PHP 8.5 hosts to Scap dsh groups - [[phab:T435393|T435393]] (andrew retrying SPF's failed log)
* 21:21 Southparkfan: switched cache-text08 backend from mediawiki14 to mediawiki16, then switched back to mediawiki14 to match Puppet state - [[phab:T435393|T435393]]
* 21:11 Southparkfan: add PHP 8.5 hosts to Scap dsh groups - [[phab:T436462|T436462]]
* 18:38 Southparkfan: decommission deployment-wikifeeds02 - [[phab:T436462|T436462]]
* 13:21 jnuche: Updating development images on contint primary for [[phab:T435931|T435931]]
* 09:16 elukey: move cx-server and citoid-beta endpoints in deployment-prep to two new Trixie VMs, update their configs and Docker images and delete the old images.
* 04:18 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1332530
=== 2026-08-27 ===
* 15:06 Krinkle: Disable duplicate publishing noise from extension-IPReputation, EIPR, [[phab:T143162|T143162]]
=== 2026-08-26 ===
* 18:07 andrewbogott: resolving rebase conflicts in /srv/git/labs/private
* 17:24 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1329617
* 14:45 hashar: Updating Castor (0.4.3..0.4.5) on all Jenkins jobs {{!}} https://gerrit.wikimedia.org/r/1329577
* 07:11 hashar: integration: updated Quibble jobs to replace deprecated `--commands` option by 1..n `--command` option(s) {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1293711 {{!}} [[phab:T321617|T321617]]
=== 2026-08-25 ===
* 21:32 brett: Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08
* 21:23 brett: revert deployment-prep acme-chief switch: acme-chief active to deployment-acme-chief05 and passive to deployment-acme-chief06
* 21:03 brett: Switch acme-chief active to deployment-acme-chief07, passive to deployment-acme-chief08
* 20:31 Southparkfan: [[phab:T401839|T401839]] - provisioned deployment-docker-wikifeeds01 (trixie) using Tofu (cloudvps-repos/deployment-prep/tofu-provisioning)
* 15:06 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/121 ([[phab:T435368|T435368]])
* 07:35 taavi@cloudcumin1001: END (PASS) - Cookbook wmcs.vps.remove_instance (exit_code=0) for instance deployment-ircd03
* 07:35 taavi@cloudcumin1001: START - Cookbook wmcs.vps.remove_instance for instance deployment-ircd03
=== 2026-08-24 ===
* 08:29 hashar: integration: add firewall rule to ssh from contint1003/contint2003 IPv6 (IPv4 was already allowed) # [[phab:T435756|T435756]]
* 08:26 hashar: integration: removing firewall rule for ssh from contint1002/contint2002 IPv4 # [[phab:T418521|T418521]]
=== 2026-08-21 ===
* 13:07 hashar: integration: triggered Doxygen doc for PersonalDashboard extension using: `zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/extensions/PersonalDashboard --change {{Gerrit|1328044}},1` # [[phab:T435392|T435392]]
=== 2026-08-20 ===
* 23:26 bd808: deployment-mediawiki14: `systemctl stop php8.3-fpm; sleep 5m; systemctl start php8.3-fpm` -- maybe the bot storm will break if we give fast 500 responses for 5 minutes.
* 14:55 James_F: jforrester@integration-castor06:$ sudo rm -rf /srv/castor/castor-mw-ext-and-skins/master/wikilambda-catalyst-end-to-end # Clear stale Catalyst castor npm downloads.
* 13:35 James_F: Zuul: [mediawiki/extensions/WikiLambda] Re-enable Catalyst
* 08:54 hashar: deployment-prep: hard reboot deployment-cache-test-08 # [[phab:T435421|T435421]]
=== 2026-08-19 ===
* 14:48 hashar: integration: updated castor save job to have rsync emit statistics {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1321049 {{!}} [[phab:T432685|T432685]]
* 12:30 James_F: Updating development images on contint primary for “fundraising: Drop libc-client-dev from the bookworm PHP 8.2 image”
=== 2026-08-18 ===
* 17:32 dancy: Rebooting deployment-mediawiki14.deployment-prep.eqiad1.wikimedia.cloud for good measure
* 17:31 dancy: rm /var/log/apache2/*.gz on deployment-mediawiki14.deployment-prep.eqiad1.wikimedia.cloud to free up ~6GB.
* 09:24 hashar: zuul: restarted zuul-web
=== 2026-08-17 ===
* 21:09 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/115
* 21:01 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/118 ([[phab:T413817|T413817]], [[phab:T407430|T407430]])
* 18:08 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/113
* 16:30 James_F: Zuul: [mediawiki/extensions/WP25EasterEggs] Archive repository, for [[phab:T418134|T418134]]
=== 2026-08-14 ===
* 18:59 James_F: Zuul: Enforce CI for mediawiki-php-<nowiki>{</nowiki>excimer,luasandbox,wikidiff2<nowiki>}</nowiki> for [[phab:T425943|T425943]]
* 18:20 James_F: Docker: [php85] Migrate to Wikimedia-provide binary, cascaded, for [[phab:T433254|T433254]]
* 14:37 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1325900
=== 2026-08-13 ===
* 19:55 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/111 ([[phab:T401115|T401115]]) (redux for typo fix)
* 19:27 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/111 ([[phab:T401115|T401115]])
* 03:05 TimStarling: created Produnto tables on beta [[phab:T421436|T421436]]
=== 2026-08-12 ===
* 12:31 hashar: integration: on Castor: `sudo rm -fR /srv/castor/*/*/mwext-phpunit-coverage*/npm` # [[phab:T427922|T427922]]
* 11:59 hashar: integration: on Castor: `sudo rm -fR /srv/castor/*/*/*codehealth*/npm` # [[phab:T427822|T427822]]
=== 2026-08-07 ===
* 16:13 hashar: integration: deleted integration-agent-[[phab:T422258|T422258]] agent # [[phab:T422258|T422258]]
* 14:31 hashar: integration: updating Quibble jobs to Quibble 1.19.0 # [[phab:T432934|T432934]] [[phab:T432943|T432943]] [[phab:T427922|T427922]]
* 07:25 hashar: Tag Quibble 1.19.0 @ {{Gerrit|a8a84ed1c0adb34688ce22066ebde8a2d584a7d4}} # [[phab:T432934|T432934]] [[phab:T432943|T432943]] [[phab:T427922|T427922]]
=== 2026-08-06 ===
* 18:02 dancy: Restarted gitlab-webhooks ([[phab:T430410|T430410]]) (revert)
* 17:55 dancy: Restarted gitlab-webhooks ([[phab:T430410|T430410]])
=== 2026-08-05 ===
* 20:31 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1184176
* 18:39 James_F: Docker: [php85] Upgrade PHP to 8.5.9
* 15:58 James_F: Zuul: [mediawiki/extensions/WikiLambda] Disable Catalyst, corrupted npm cache
=== 2026-08-04 ===
* 19:06 dancy: Updating buildkitd to v0.32.2 on gitlab-cloud-runners (staging and production) ([[phab:T433985|T433985]])
* 16:48 James_F: Docker: [php83] Upgrade PHP to 8.3.33
=== 2026-08-03 ===
* 20:45 dancy: Updating buildkitd to v0.32.1 on gitlab-cloud-runners (staging and production) ([[phab:T433879|T433879]])
* 16:40 hashar: gerrit: added Vaughn Walters to integration group until he get added to the ciadmin LDAP group {{!}} [[phab:T433615|T433615]]
* 16:37 James_F: Zuul: Drop REL1_44 testing, EOL, for [[phab:T428911|T428911]]
=== 2026-07-31 ===
* 18:14 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/114
* 12:23 hashar: integration: sudo cumin --force -p 0 'name:docker' 'rm -fR /srv/jenkins/workspace/*pipeline*'
=== 2026-07-30 ===
* 17:28 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover/mediawiki-libs-node-cssjanus/ # [[phab:T424419|T424419]]
=== 2026-07-29 ===
* 21:33 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/109
* 18:33 dancy: Buildkit v0.32.0 deployed to gitlab-cloud-runners staging and production ([[phab:T433520|T433520]])
=== 2026-07-24 ===
* 17:59 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1316011
=== 2026-07-23 ===
* 15:06 dancy: Deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314124 ([[phab:T295351|T295351]])
=== 2026-07-22 ===
* 21:32 dancy: Moved /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/npm to /srv/castor-debug-[[phab:T20260722|T20260722]]-npm-torn-cacache/ on integration-castor06
* 18:41 dancy: Zuul dependencies upgraded and Zuul restarted.
* 18:25 dancy: Zuul is currently broken due to the Gerrit SSH key update. I'm investigating
* 17:24 brennen: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1314014/1 ([[phab:T432886|T432886]])
* 16:58 dancy: Restarting Gerrit ([[phab:T240266|T240266]]) (again)
* 15:59 dancy: Restarting Gerrit ([[phab:T240266|T240266]])
* 09:54 James_F: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure for {{Gerrit|1313321}}, per Lucas_WMDE.
=== 2026-07-21 ===
* 15:59 dancy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/c/integration/config/+/1313220
* 07:08 hashar: integration: cleaned up Jenkins workspace on integration-agent-docker-1084
=== 2026-07-20 ===
* 19:36 dancy: Restarting jenkins on contint1003 to clear out hours-long stuck jobs
=== 2026-07-16 ===
* 23:07 mutante: gerrit1003/2002/2003: rm /srv/gerrit/.ssh/config and revert gerrit:1311577 which fixed [[phab:T398401|T398401]] but caused [[phab:T432413|T432413]] - replication is working again
* 18:51 dancy: Updated buildkit to v0.31.2 gitlab-cloud-runners (staging and production) ([[phab:T432360|T432360]])
* 18:37 dancy: Restarted Jenkins to unstick jobs.
* 18:33 dancy: Jenkins placed in shutdown mode in preparation for a restart
* 17:07 bd808: Hard reboot of deployment-cache-text08.deployment-prep.eqiad1.wikimedia.cloud via Horizon; console shows OOM ([[phab:T432374|T432374]])
* 16:40 brennen: contint1002: chown and chmod on /srv/zuul/git/mediawiki/extensions/CampaignEvents/.git per https://www.mediawiki.org/wiki/Continuous_integration/Zuul#Very_high_queue_of_merger:merge_functions
* 16:34 dancy: Restarted docker on contint2003 ([[phab:T432326|T432326]])
* 16:33 dancy: Restarted docker on contint1003 ([[phab:T432326|T432326]])
* 16:32 dancy: Restarted docker on contint1003
* 12:31 Krinkle: krinkle@contint1003:~$ sudo /usr/sbin/service jenkins restart
* 12:10 Krinkle: krinkle@contint1002: zuul restart
=== 2026-07-15 ===
* 18:58 James_F: dev-images: Re-build PHP images for latest point releases, for [[phab:T431099|T431099]]
* 09:13 James_F: Docker: [php83] Update PHP to 8.3.32 for [[phab:T431099|T431099]]
=== 2026-07-14 ===
* 11:53 James_F: Docker: [commit-message-validator] Update to v3.0.0, for [[phab:T431799|T431799]]
* 08:37 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud
=== 2026-07-13 ===
* 15:03 dancy: Updating gitlab-cloud-runners to v19.0.2
=== 2026-07-10 ===
* 11:56 James_F: Zuul: Disable all browser tests on release branches except Wikibase's, for [[phab:T430415|T430415]]
* 09:16 James_F: Docker: [commit-message-validator] Update to 2.3.0
=== 2026-07-09 ===
* 11:55 hashar: retriggering postmerge change for [[phab:T431582|T431582]]: zuul enqueue --trigger gerrit --pipeline postmerge --project machinelearning/liftwing/inference-services --change {{Gerrit|1308631}},3
=== 2026-07-08 ===
* 20:58 brennen: patchdemo: deployed https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/367 ([[phab:T427964|T427964]])
* 19:21 mutante: gerrit - replacing registerEmailPrivateKey in Gerrit config - this invalidates pending/outstanding email validation links for gerrit users - but does not affect active accounts or already verified email addresses
* 14:36 hashar: contint1003, contint2003: manually installed `docker-buildx` Debian package to validate https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308659 # [[phab:T431582|T431582]]
* 13:04 hashar: deployment-prep: git repack on /srv/mediawiki-staging/php-master
* 12:27 hashar: deployment-prep: on deployment server: clearing old branches for mediawiki/extensions and mediawiki/skins # [[phab:T428864|T428864]]
=== 2026-07-07 ===
* 19:41 mutante: contint1003/2003 - add jenkins-agent user to docker group; restart jenkins
* 17:17 hashar: gerrit: deleted /srv/gerrit/java_pid3571660.hprof
=== 2026-07-06 ===
* 20:06 dancy: Updated buildkitd to v0.31.1 in gitlab-cloud-runners ([[phab:T429988|T429988]])
* 08:11 James_F: Zuul: Add WikimediaAntiAbuse extension, for [[phab:T431023|T431023]]
=== 2026-07-03 ===
* 15:25 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add Elastica dep too
* 15:09 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CirrusSearch dep
* 11:22 James_F: Zuul: Make the in-mediawiki-tarball template real
* 10:06 James_F: Zuul: [mediawiki/extensions/TestKitchen] Don't drop from release branches
* 09:53 James_F: Docker: [quibble-coverage] Update phpunit-patch-coverage to 0.0.18, for [[phab:T423987|T423987]] and [[phab:T425807|T425807]]
=== 2026-07-02 ===
* 13:27 James_F: Docker: [composer-scratch] Upgrade composer to 2.10.2 and cascade, for [[phab:T428570|T428570]]
* 10:47 hashar: zuul1002: running Puppet agent to drop `wikimediacloud.org` from `no_proxy` {{!}} [[phab:T430479|T430479]]
=== 2026-07-01 ===
* 13:39 hashar: integration: added timestamping to operations-puppet-catalog-compiler and operations-puppet-catalog-compiler-puppet7-test jobs
* 02:45 hashar: gerrit: on gerrit2003 deleted /srv/gerrit/java_pid3520115.hprof (the JVM apparently died at some point, I assume due to heavy crawling)
=== 2026-06-30 ===
* 15:13 dancy: Rebooting deployment-mwlog02.deployment-prep to clear stuck udp2log processes
=== 2026-06-29 ===
* 09:32 hashar: gerrit: deleted repository phabricator/extensions/BurnDownCharts , created in July 2014, had no commit/changes
=== 2026-06-28 ===
* 15:30 hashar: Updated integration/zuul-jobs from upstream (c75fe6ef19c..fc4af6d4471), notably to remove `requestsexceptions` in `upload-logs-swift` role # [[phab:T430458|T430458]]
=== 2026-06-26 ===
* 13:59 Krinkle: [[phab:T429658|T429658]] krinkle@doc1004:/srv/doc/cover-extensions$ sudo -u doc-uploader rm -rf ShortUrl/
* 13:59 Krinkle: [[phab:T429658|T429658]] krinkle@doc2003:/srv/doc/cover-extensions$ sudo -u doc-uploader rm -rf ShortUrl/
=== 2026-06-25 ===
* 20:57 dancy: Restarting Jenkins to unstick builds
* 20:49 dancy: Investigating castor-save-workspace-cache clog
* 16:50 inflatador: add 60GB cinder vol to deployment-cirrussearch15 [[phab:T425585|T425585]]
* 14:25 inflatador: delete unused servers deployment-cirrussearch1[2-4] [[phab:T425585|T425585]]
* 10:07 James_F: Docker: Bump Node 24 / Node 26 to new releases
=== 2026-06-24 ===
* 17:26 dancy: Set `profile::puppetserver::autosign: /usr/local/sbin/validatecloudvpsfqdn.py` in hiera config for deployment-puppetserver prefix ([[phab:T429413|T429413]])
* 15:12 dancy: sudo systemctl restart php8.3-fpm on deployment-jobrunner05 (Attempting to resolve logspam)
=== 2026-06-23 ===
* 14:31 brennen: deploying patchdemo for https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/361
=== 2026-06-22 ===
* 23:19 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/mediawiki-core/master/mediawiki-node24/ #[[phab:T429824|T429824]]
* 23:07 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ #[[phab:T429824|T429824]] (again)
* 22:35 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ #[[phab:T429824|T429824]]
* 19:30 thcipriani: thcipriani@integration-castor06:~$ sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-with-gated-extensions-vendor-mysql-php83 #[[phab:T429824|T429824]]
=== 2026-06-19 ===
* 20:13 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1304632
* 09:39 Lucas_WMDE: deployment-deploy04: used createAndPromote to restore User:Lucas Werkmeister (WMDE) to bureaucrat (dewiki, enwiki, metawiki, wikidatawiki) and wikidata-staff (wikidatawiki) after groups were very unhelpfully removed, presumably due to 2FA enforcement, without apparent warning or announcement of any kind
* 07:07 dcausse: restarted php-fpm on deployment-jobrunner05 (inconsistent php state: MediaWiki\Extension\PageAssessments\HookHandler\ParserHooks::__construct(): Argument #3 ($config) must be of type MediaWiki\Config\Config, MediaWiki\Page\WikiPageFactory given)
* 06:13 hashar: Upgrading Quibble jobs to 1.18.2 (retry git connections on reset) https://gerrit.wikimedia.org/r/c/integration/config/+/1304181 # [[phab:T420865|T420865]]
=== 2026-06-18 ===
* 19:52 dancy: Building quibble 1.18.2 images on contint primary ([[phab:T420865|T420865]])
* 18:50 dancy: Tag Quibble 1.18.2 @ {{Gerrit|152b497d317eaafc3cec334d5ce7e549697a2980}} # [[phab:T420865|T420865]]
* 15:15 dancy: Deleting deployment-db11 and deployment-db14 ([[phab:T428910|T428910]])
* 07:23 dcausse: reindexing all wikis to opensearch2 ([[phab:T425585|T425585]], [[phab:T427196|T427196]])
=== 2026-06-17 ===
* 23:16 dduvall: restarted zuul to clear up 5 hrs of stuck queues
* 23:03 mutante: re-enabled puppet on contint1003 - triple checked puppet does NOT start jenkins anymore. BOTH masked AND stopped while running untouched on contint1002. after: gerrit:1303578 {{!}} ([[phab:T418521|T418521]]) ([[phab:T428791|T428791]])
* 18:51 dduvall: performed stop/start of jenkins service on contint1002 following failed safeRestart
* 18:45 dduvall: restarting jenkins due to stuck zuul queues
* 14:56 inflatador: mwscript /srv/mediawiki-staging/php-master/extensions/CirrusSearch/maintenance/ForceSearchIndex.php --wiki=enwikibooks
=== 2026-06-16 ===
* 16:00 dancy: sudo keyholder arm on deployment-deploy04.deployment-prep
* 15:48 dancy: Resizing deployment-deploy04.deployment-prep from g4.cores4.ram8.disk20 to g4.cores8.ram16.disk20 ([[phab:T429364|T429364]])
* 15:31 dancy: Turning off deployment-db11 and deployment-db14
=== 2026-06-15 ===
* 23:40 dancy: systemctl restart php8.3-fpm on deployment-jobrunner05 to reload beta db configuration ([[phab:T428930|T428930]])
* 21:00 dancy: deployment-db15 promoted to master ([[phab:T428930|T428930]])
* 20:53 dancy: deployment-db11.deployment-prep going read-only
* 20:48 dancy: deployment-db11.deployment-prep will be going read-only soon while db15 is being promoted to primary.
* 20:26 dancy: Added deployment-db16 ([[phab:T429245|T429245]])
* 18:26 dancy: Rebooting deployment-jobrunner05
=== 2026-06-12 ===
* 21:45 dancy: Unstuck wmf-beta-update-all service on deployment-deploy04.deployment-prep (sudo systemctl stop wmf-beta-update-all)
* 18:00 thcipriani: unmasking jenkins on contint1002 and restarting
* 17:49 thcipriani: attempting to cancel castor-save-workspace-cache {{Gerrit|6710545}}
* 15:19 James_F: Docker: [php83] Re-platform to Debian Bookworm, for [[phab:T383337|T383337]]
* 15:07 dancy: deployment-db15 configured as a replica of deployment-db11 ([[phab:T428930|T428930]])
* 10:21 Krinkle: `krinkle@<nowiki>{</nowiki>doc1004,doc2003<nowiki>}</nowiki>:/srv/doc/mediawiki-core$ sudo -u doc-uploader rm -rf list/` - remove doc build for git-tag test.
=== 2026-06-11 ===
* 14:46 hashar: for minor in $(seq 21 42); do ./.tox/make-release/bin/python -u ./make-release/branch.py --delete --abandon --bundle '*' "REL1_$minor"; done;
* 14:46 hashar: On all MediaWiki repos, converting old release branches up to REL1_42 included to tags. Last time I missed non wmf repo # [[phab:T380841|T380841]] {{!}} [[phab:T428864|T428864]]
* 09:45 hashar: Converted mediawiki/core branches REL1_39, REL1_40, REL1_41, REL1_42 to tags # [[phab:T428864|T428864]]
* 09:18 hashar: Converting REL1_42 branches to tags # [[phab:T428864|T428864]]
* 09:18 hashar: Converting REL1_41 branches to tags # [[phab:T428864|T428864]]
* 09:05 hashar: Converting REL1_40 branches to tags # [[phab:T428864|T428864]]
* 08:56 hashar: Converting REL1_39 branches to tags # [[phab:T428864|T428864]]
* 08:40 hashar: gerrit: deleted mediawiki/core branch "development" that pointed to {{Gerrit|7f622781cc31053b121f6f4ddbff506cba10d38e}} (which is contained by master). Had probably been created by a direct push.
=== 2026-06-10 ===
* 16:06 James_F: Docker: Provide quibble-bookworm, for [[phab:T362705|T362705]]
=== 2026-06-04 ===
* 23:39 jeena: Updating development images on contint primary for [[phab:T424691|T424691]]
* 12:30 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc # fix failure seen in mwext-node24-rundoc 4812
* 08:35 hashar: Built Docker images `docker-registry.wikimedia.org/releng/java21:0.1` and `docker-registry.wikimedia.org/releng/maven-java21:0.1` # [[phab:T412978|T412978]]
=== 2026-06-03 ===
* 14:19 jnuche: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/107
* 12:54 James_F: Zuul: Add Rae 5e as a trusted user
* 08:26 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1296559 "inference-services: Add LLM generated editing suggestions CI/CD pipelines." # [[phab:T427794|T427794]]
=== 2026-06-02 ===
* 23:23 thcipriani: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/1296695
* 23:10 thcipriani: tag quibble 1.18.1 @ {{Gerrit|4b7959553c095d811426f394165c69ecc13a44eb}}
* 18:13 brennen: devtools phab/phorge: deployed work/2026-06-01-merge-phorge to https://phabricator.wmcloud.org/ for testing ([[phab:T410849|T410849]])
* 16:05 hashar: integration: upgraded pypy from 7.3.11 (py3.9) to 7.3.20 (py3.11) # [[phab:T423607|T423607]]
* 15:49 jnuche: Updating buildkitd to v0.30.0 in gitlab-cloud-runners ([[phab:T426212|T426212]])
* 15:24 hashar: Building docker images for https://gerrit.wikimedia.org/r/c/integration/config/+/1295865 # [[phab:T423607|T423607]]
* 14:57 jnuche: Jenkins/Zuul is back
* 14:38 jnuche: restarting Jenkins
* 14:28 jnuche: bring back castor node, that didn't help
* 14:23 jnuche: trying to reconnect castor node, see if that helps somehow
* 14:12 jnuche: option to "Enable Gearman" times out. Can't re-enable from UI. Gearman plugin logs are empty. Neat
* 14:05 jnuche: trying to reconnect Gearman
=== 2026-06-01 ===
* 21:34 jeena: Updating development images on contint primary for [[phab:T424663|T424663]]
* 10:58 hashar: gerrit: flushed `ldap_usernames` cache in case a missing account ended up being cached there # [[phab:T427792|T427792]]
=== 2026-05-28 ===
* 16:48 hashar: castor: nuked SonarQube cache: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-codehealth-master-non-voting/sonar/ # [[phab:T427471|T427471]]
* 07:05 hashar: integration: delete deployment-deploy04 agent from Jenkins controller. Batch job has been migrated to a systemd driven script # [[phab:T256168|T256168]]
* 06:29 hashar: integration: delete integration-castor05 agent from Jenkins, replaced by integration-castor06 # [[phab:T421114|T421114]]
=== 2026-05-27 ===
* 21:07 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1294425
* 09:49 hashar: Updated helm-lint Jenkins job to use releng/helm-linter:0.8.0 image # [[phab:T424824|T424824]]
* 08:25 codders: integration-castor06: sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/
=== 2026-05-26 ===
* 15:26 James_F: Zuul: Add PHP 8.5 to a few missing PHP pipelines, oops
* 14:46 hashar: Updating Quibble Jenkins jobs to stop Supervisord from spawning Memcached # [[phab:T397810|T397810]]
* 14:07 hashar: Updated integration-quibble-* jobs in order to validate running tests with Supervisord not managing memcached # [[phab:T397810|T397810]]
* 13:15 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-docs-publish # fix failure seen in mwext-node24-docs-publish 981, 985, 987
* 09:05 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1292620 (Introduce Phan composer job - [[phab:T231966|T231966]])
=== 2026-05-21 ===
* 16:33 hashar: Reloaded Zuul to enable Node24 CI job for `labs/tools/wdaudiolex-fe` # [[phab:T426366|T426366]]
=== 2026-05-19 ===
* 18:07 James_F: Zuul: [mediawiki/libs/ZestJQ] Add basic PHP and Node CI
=== 2026-05-18 ===
* 21:12 James_F: Zuul: [operations/software/gerrit] Add Node26 as experimental
* 21:10 mutante: gerrit-replica.wikimedia.org, gerrit-spare.wikimedia.org - rebooting backends
* 20:57 James_F: Zuul: [integration/docroot] Test in PHP 8.3+, dropping 8.2
* 20:56 James_F: Zuul: [analytics/wmde/scripts] Test in PHP 8.3+, dropping 8.
* 20:18 James_F: Zuul: [wikimedia/fundraising/dash] Replace Node 20 testing with Node 24
* 20:18 James_F: Zuul: Migrate various labs things to Node 24
* 20:02 James_F: Docker: [ajv, sonar-scanner] Migrate to Node 24
* 19:58 James_F: Zuul: Migrate various production/CI things to Node 24
* 18:17 mutante: releases.wikimedia.org - rebooting backends
* 18:14 mutante: rebooting production gitlab-runners
* 18:12 dancy: gitlab-cloud-runners have been revived.
* 18:11 James_F: Zuul: [design/codex] Switch CI to Node 24
* 15:52 dancy: gitlab-cloud-runners are in a broken state. I'm investigating
* 14:27 hashar: Upgrading Quibble jobs to 1.18.0
* 09:29 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure seen in mediawiki-node24 22260
=== 2026-05-15 ===
* 18:36 dancy: Upgraded gitlab-cloud-runners (prod) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 18:24 dancy: Upgrading gitlab-cloud-runners (prod) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 18:11 dancy: Upgraded gitlab-cloud-runners (staging) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 17:59 dancy: Upgrading gitlab-cloud-runners (staging) from 1.35.1-do.5 to 1.35.1-do.6 ([[phab:T426436|T426436]])
* 13:02 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1287828 [[phab:T426392|T426392]]
=== 2026-05-13 ===
* 12:42 James_F: Zuul: [mediawiki/extensions/Springboard] Add AdminLinks Phan dependency
* 12:42 James_F: Zuul: [mediawiki/extensions/ChatBot] Add dependencies on VisualEditor and BlueSpiceFoundation
* 12:42 James_F: Zuul: [mediawiki/extensions/ChatIntegration] Add dependency on VisualEditor
* 12:37 James_F: Zuul: [mediawiki/extensions/WikiLambda] Drop AF and SB deps down to phan-only, for [[phab:T423180|T423180]]
=== 2026-05-12 ===
* 20:57 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/104 ([[phab:T424774|T424774]])
* 18:08 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add AF and SB deps for [[phab:T423180|T423180]]
* 14:18 atsukoito: PrivateSettings: empty $wgOpensearchCredentials for opensearch-on-k8s synced to deploy04 by Reedy
* 13:04 atsukoito: PrivateSettings: credentials for opensearch-on-k8s ttmserver-test
* 11:50 James_F: Zuul: [machinelearning/liftwing/inference-services] Add qwen36 llm model CI/CD pipelines, for [[phab:T425680|T425680]]
* 11:46 James_F: Zuul: Add experimental php-pie-build* jobs to other PHP extensions, for [[phab:T425943|T425943]]
* 11:37 James_F: Zuul: [mediawiki/php/wikidiff2] Add experimental php-pie-build* jobs, for [[phab:T425943|T425943]]
* 10:05 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-with-Wikibase-extensions-browser-tests-only-vendor-php83 # fix failure seen in quibble-with-Wikibase-extensions-browser-tests-only-vendor-php83 7817
* 08:44 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # broken Cypress cache? hopefully fix failure seen in quibble-vendor-mysql-php83-selenium 51633
=== 2026-05-11 ===
* 18:28 James_F: Docker: Add changes to php-compile images for PIE, for [[phab:T425943|T425943]]
* 16:06 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/ # broken Cypress cache? hopefully fix failure seen in quibble-vendor-mysql-php83-selenium 51439 and 51452
=== 2026-05-09 ===
* 20:46 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1285498
=== 2026-05-07 ===
* 22:53 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/105
=== 2026-05-06 ===
* 18:13 bd808: Unblock 88.165.192.0/19
* 18:03 bd808: Unblock 94.208.0.0/14
* 17:56 bd808: Unblock 84.226.0.0/16
* 17:41 bd808: Unblock 94.34.0.0/16
* 17:35 bd808: Unblock 109.134.0.0/16
=== 2026-05-05 ===
* 21:20 James_F: Zuul: Provide Node 26 experimental jobs everywhere needed
* 21:04 James_F: Docker: Provide initial Node 26 images
* 19:01 James_F: Zuul: [mediawiki/extensions/PageAssessments] Add Scribunto dependency, for [[phab:T396135|T396135]]
* 14:58 dancy: rm /var/log/<nowiki>{</nowiki>user.log.1,syslog.1,messages.1<nowiki>}</nowiki> on deployment-eventgate-4.deployment- prep ([[phab:T425429|T425429]])
=== 2026-05-04 ===
* 15:19 dancy: Upgrading gitlab cloud runners (prod) from 1.35.1-do.3 to 1.35.1-do.5
* 14:51 dancy: Upgrading gitlab cloud runners (staging) from 1.35.1-do.3 to 1.35.1-do.5
* 10:40 James_F: Zuul: Provide non-voting PHP 8.4/8.5 Quibble jobs for bluespice template
=== 2026-05-02 ===
* 20:49 James_F: Zuul: [mediawiki/core] Enforce PHP 8.4 & 8.5 on release branches, all pass
* 19:27 James_F: Zuul: Provide non-voting PHP 8.4/8.5 Quibble jobs for MW release branches
* 19:19 James_F: Zuul: [mediawiki/extensions/BlogPage] Add dependencies
* 16:48 James_F: Hard-restarting Zuul to clear the huge number of i18n updates being re-submitted.
* 15:48 James_F: Zuul: [wikimedia-cz/*] Test in PHP 8.3+, dropping 8.2
* 14:02 TheresNoTime: Add bvibber to deployment-prep project
* 09:08 James_F: Docker: [quibble-*] Add php-luasandbox so we can test both modes in Scribunto
=== 2026-05-01 ===
* 15:42 James_F: Zuul: [wikimedia/lucene-explain-parser] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: Zuul: [wikimedia/textcat] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: Zuul: [mediawiki/tools/ParseWiki] Test in PHP 8.3+, dropping 8.2
* 15:42 James_F: zuul: Add ToprakM to CI allowlist
* 15:19 James_F: Zuul: [translatewiki] Test in PHP 8.3+, dropping 8.2
* 15:10 James_F: Zuul: [mediawiki/extensions/WikiEditor] Add TestKitchen as a dependency, for [[phab:T425076|T425076]]
* 12:40 James_F: Zuul: [mediawiki/tools/code-utils] Test in PHP 8.3+, dropping 8.2
* 08:02 James_F: Zuul: Update xtex's e-mail in the allowlist
* 07:37 James_F: Zuul: Switch release branches' selenium jobs to PHP 8.3
* 07:33 James_F: Zuul: Test Wikimedia production libraries in PHP 8.3+, dropping 8.2
=== 2026-04-30 ===
* 21:36 brennen: gitlab-webhooks: building & restarting to deploy https://gitlab.wikimedia.org/repos/releng/gitlab-webhooks/-/merge_requests/40
* 20:26 James_F: Zuul: [mediawiki/tools/api-testing] Make PHP 8.5 CI voting
* 20:16 James_F: jforrester@doc1004:~$ # sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/WebAuthn/ # [[phab:T415832|T415832]]
* 20:14 James_F: Zuul: [mediawiki/extensions/WebAuthn] Archive, for [[phab:T415832|T415832]] / [[phab:T303495|T303495]]
* 17:16 brennen: wikibugs: most maintainers at hackathon, so go release-engineering added as a maintainer while looking to debug error at https://gitlab.wikimedia.org/toolforge-repos/wikibugs2/-/jobs/810904
* 15:19 mutante: upgrading zuul to 14.2.0-1 on "new zuul" machines ([[phab:T424879|T424879]])
=== 2026-04-29 ===
* 15:49 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add ConfirmEdit dependency, for [[phab:T424597|T424597]]
* 15:36 James_F: Zuul: Drop experimental node22 jobs, never used in practice
* 15:28 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1279392, https://gerrit.wikimedia.org/r/1279397
=== 2026-04-28 ===
* 18:11 bd808: Unblock 86.0.0.0/16
* 17:41 bd808: Unblock 79.192.0.0/10
* 17:07 James_F: Zuul: [mediawiki/tools/phpunit-patch-coverage] Drop PHP 8.2 testing
* 17:07 James_F: Zuul: [mediawiki/tools/minus-x] Drop PHP 8.2 testing
* 17:07 James_F: Zuul: [mediawiki/tools/codesniffer] Drop PHP 8.2 testing
* 16:32 James_F: Zuul: [mediawiki/services/jobrunner] Drop PHP 8.2 testing
* 13:34 James_F: Zuul: [mediawiki/tools/phan] Drop PHP 8.2 testing
* 13:34 James_F: Zuul: [oojs/ui] Drop PHP 8.2 testing
* 13:14 James_F: Zuul: [mediawiki/tools/phan/SecurityCheckPlugin] Drop PHP 8.2 CI
* 10:40 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-rundoc/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud to fix failure seen in mwext-node24-rundoc #1717
* 00:03 bd808: Increase parallelism for wmf-beta-update-databases.py ([[phab:T256168|T256168]])
=== 2026-04-27 ===
* 22:11 bd808: Beta Cluster MediaWiki update logs now available via https://beta-update.wmcloud.org/ ([[phab:T256168|T256168]])
* 21:57 bd808: Add web security group to deployment-deploy04 ([[phab:T256168|T256168]])
* 20:45 James_F: Zuul: Restrict mw*-codehealth-patch jobs to master only, for [[phab:T424573|T424573]]
* 17:16 James_F: Docker: [mediawiki-phan-taint-check-demo] Re-platform to Trixie and so PHP 8.4
* 15:53 James_F: Zuul: [mediawiki/extensions/ReportIncident] Add TestKitchen phan dependency, for [[phab:T424220|T424220]]
* 14:32 James_F: Zuul: Drop PHP 8.2 enforcement from MediaWiki things for master and REL1_46 for [[phab:T358667|T358667]]
* 12:38 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mwext-node24-docs-publish # fix failure seen in mwext-node24-docs-publish 383
* 09:18 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover/mediawiki-libs-node-cssjanus/ # [[phab:T424419|T424419]]
=== 2026-04-26 ===
* 20:49 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1276777
=== 2026-04-24 ===
* 22:48 dduvall: merged zuul3 branch of integration/config into master and pushed (in preparation for https://gerrit.wikimedia.org/r/c/operations/puppet/+/1277198)
* 12:27 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1276428
=== 2026-04-23 ===
* 23:57 bd808: Set `profile::beta::autoupdater::run_updater: true` for deployment-deploy04 via Horizon ([[phab:T256168|T256168]])
* 22:58 bd808: bd808@deployment-deploy04 `sudo -u jenkins-deploy /usr/local/bin/wmf-beta-update-all`
* 22:36 bd808: bd808@deployment-deploy04 `sudo -u mwdeploy /usr/local/bin/wmf-beta-update-all`
* 22:16 bd808: Disabled https://integration.wikimedia.org/ci/view/Beta/job/beta-update-databases-eqiad so that replacement script can be tested ([[phab:T256168|T256168]])
* 22:12 bd808: Disabled https://integration.wikimedia.org/ci/job/beta-code-update-eqiad so that replacement script can be tested ([[phab:T256168|T256168]])
* 22:02 bd808: Cherry-picked {{gerrit|1276813}} to deployment-puppetserver-1 ([[phab:T256168|T256168]])
* 20:11 James_F: Zuul: [wikibase/*] Replace CI testing in Node 20 with Node 24
* 20:11 James_F: Zuul: [wikidata/query/*] Replace CI testing in Node 20 with Node 24
* 20:11 James_F: Zuul: [analytics/*] Replace CI testing in Node 20 with Node 24
* 20:10 James_F: Zuul: [mediawiki/tools/*] Replace CI testing in Node 20 with Node 24
* 20:06 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.34.5-do.3 to 1.35.1-do.3 ([[phab:T423726|T423726]])
* 19:55 James_F: Zuul: [jquery-client] Replace CI testing in Node 20 with Node 24
* 19:51 James_F: Zuul: [wikipeg] Drop testing in Node 20 and Node 22
* 19:47 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.34.5-do.3 to 1.35.1-do.3 ([[phab:T423726|T423726]])
* 19:37 James_F: Zuul: [oojs/ui] Drop CI testing in Node 20 and Node 22
* 19:37 James_F: Zuul: [oojs/js] Drop CI testing in Node 20 and Node 22
* 19:37 James_F: Zuul: [unicodejs] Replace CI testing in Node 20 with Node 24
* 19:36 James_F: Zuul: [wikimedia/portals] Drop CI testing in Node 20 and Node 22
* 18:57 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.33.9-do.3 to 1.34.5-do.3 ([[phab:T423726|T423726]])
* 18:39 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.33.9-do.3 to 1.34.5-do.3 ([[phab:T423726|T423726]])
* 18:18 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.33.9-do.2 to 1.33.9-do.3 ([[phab:T423726|T423726]])
* 17:58 James_F: Zuul: [mediawiki/extensions/OAuth] Add dependency on CentralAuth, for [[phab:T415281|T415281]]
* 17:56 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.32.13-do.2 to 1.33.9-do.3 ([[phab:T423726|T423726]])
* 16:35 James_F: Zuul: Enforce PHP 8.5 CI for MW things in master (and REL1_46), for [[phab:T411814|T411814]]
* 16:19 James_F: Zuul: [mediawiki/services/parsoid] Enable PHP 8.5 CI
* 15:47 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add AntiSpoof dependency, for [[phab:T420548|T420548]]
* 14:20 Lucas_WMDE: ssh integration-castor06.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24 # fix failure seen in mediawiki-node24 8385 and 8405
* 12:56 James_F: Zuul: [mediawiki/extensions/GrowthExperiments] Add CentralNotice dependency, for [[phab:T422082|T422082]]
=== 2026-04-22 ===
* 00:07 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add MF dependency, for [[phab:T424113|T424113]]
=== 2026-04-21 ===
* 23:26 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CommunityConfiguration dep too, for [[phab:T394410|T394410]]
* 23:17 James_F: Zuul: [mediawiki/extensions/DiscussionTools] Add standalone test jobs, for [[phab:T422031|T422031]]
* 20:47 inflatador: updating cirrussearch hosts to Trixie/OpenSearch 2 [[phab:T421763|T421763]]
* 20:38 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add CommunityConfiguration phan dep, for [[phab:T394410|T394410]]
* 20:17 bd808: Running tofu for [[phab:T421244|T421244]]
* 18:00 James_F: Zuul: [mediawiki/extensions/WatchAnalytics] Add ApprovedRevs Phan dependency
* 16:35 bd808: Unblock 79.116.0.0/16
* 13:34 James_F: Zuul: [mediawiki/extensions/WikiLambda] Add TestKitchen phan dep, for [[phab:T415254|T415254]]
* 13:27 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add CentralAuth dependency, for [[phab:T420548|T420548]]
=== 2026-04-20 ===
* 23:56 bd808: Unblock 76.157.0.0/16
* 18:28 dancy: Upgrading gitlab cloud runners (staging) to 1.33.9-do.2 ([[phab:T423726|T423726]])
* 18:28 dancy: Upgrading gitlab cloud runners (staging) ([[phab:T423726|T423726]])
* 18:19 James_F: jjb: All 486 (!) jobs now updated for [[phab:T423622|T423622]]
* 18:18 bd808: Unblock 113.128.0.0/15
* 15:03 James_F: Docker: Bump ci-bullseye/-bookworm/-trixie for mirrors.wm.org removal, [[phab:T423622|T423622]]
=== 2026-04-19 ===
* 19:53 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1272752
=== 2026-04-17 ===
* 21:07 thcipriani: marking integration-agent-1080 offline for experimentation
* 19:30 thcipriani: reconfiguring castor-save-workspace-cache with https://gerrit.wikimedia.org/r/1273935
* 17:47 dancy: Upgrading gitlab cloud runners (prod) k8s from 1.32.10-do.1 to 1.32.13-do.2 ([[phab:T423726|T423726]])
* 16:49 dancy: Upgrading gitlab cloud runners (staging) k8s from 1.32.10-do.1 to 1.32.13-do.2 ([[phab:T423726|T423726]])
=== 2026-04-16 ===
* 20:49 dduvall: creating integration/zuul-jobs repo to serve as a mirror of opendev.org/zuul/zuul-jobs ([[phab:T406384|T406384]])
* 13:38 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1272711 [[phab:T423568|T423568]]
* 11:07 Silvan_WMDE: sudo -u jenkins-deploy rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node24/ # run on integration-castor06.integration.eqiad1.wikimedia.cloud
=== 2026-04-15 ===
* 20:05 James_F: Zuul: Configure REL1_46 CI, for [[phab:T423257|T423257]]
* 17:44 bd808: Unblock 176.0.0.0/13
* 17:39 bd808: Unblock 46.128.0.0/16
* 17:32 bd808: Unblock 176.86.0.0/16
* 16:39 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/commit/127d783b2176ac60b646a5fa4f1b1a872ca66340
* 15:33 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/100
* 01:02 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/99
=== 2026-04-14 ===
* 20:42 James_F: Docker: [composer-scratch] Upgrade composer to 2.9.7 and cascade
* 16:35 bd808: Unblock 88.112.0.0/14
* 00:48 bd808: Unblock 24.6.0.0/16
* 00:42 bd808: Unblock 152.231.48.0/20
=== 2026-04-13 ===
* 22:00 James_F: Zuul: [mediawiki/vendor] Drop accidental Wikibase browser tests on branches
* 20:28 James_F: Zuul: [mediawiki/extensions/Chart] Drop Doxygen publish job, not used
* 14:42 James_F: Zuul: [mediawiki/extensions/WikimediaCustomizations] Add FlaggedRevs dep, for [[phab:T421011|T421011]]
=== 2026-04-12 ===
* 18:21 James_F: jforrester@contint1002:~$ sudo /usr/sbin/service zuul restart && tail -f -n100 /var/log/zuul/zuul.log # [[phab:T423027|T423027]]
=== 2026-04-10 ===
* 23:22 James_F: jforrester@contint1002:~$ zuul enqueue --trigger gerrit --pipeline postmerge --project mediawiki/extensions/ReadingLists --change {{Gerrit|1269498}},2 # [[phab:T422976|T422976]]
* 23:20 James_F: Zuul: [mediawiki/extensions/ReadingLists] Publish JS coverage, for [[phab:T422976|T422976]]
* 23:13 James_F: Zuul: Migrate a few straggler Node 20 MediaWiki things to Node 24
* 23:01 James_F: Zuul: Move all MediaWiki things from mediawiki-node20 to mediawiki-node24
* 21:59 James_F: Docker: Bump Node base images to March releases and cascade; Upgrade Quibble images from Node 20 to Node 24
* 10:24 hashar: Updating all Quibble jobs to 1.17.1
* 10:22 hashar: Updated PostgreSQL jobs to Quibble 1.17.1 # [[phab:T422110|T422110]]
* 10:22 hashar: Updated apitesting job to Quibble 1.17.1 # [[phab:T422843|T422843]] [[phab:T418743|T418743]]
* 09:51 hashar: Tag Quibble 1.17.1 @ {{Gerrit|0a1ab3b7c3dfee36c9bc2e9b049957d94e190e85}}
=== 2026-04-09 ===
* 15:13 hashar: Rolling back Quibble jobs to 1.16.0 (api-testing stage fails due to missing npm install step`
* 14:58 hashar: Upgrading Quibble jobs to 1.17.0
* 14:23 hashar: Tagged Quibble 1.17.0 @ {{Gerrit|864381c6b63bdbcd8c74a3162c406fffcaaf8694}}
* 07:48 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1268559 "Zuul: use standalone jobs for GrowthExperiments Cypress tests" {{!}} [[phab:T417412|T417412]]
=== 2026-04-08 ===
* 22:19 dancy: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1269068
* 22:01 bd808: Unblock 95.216.12.170/32 ([[phab:T422751|T422751]])
* 19:26 brennen: gitlab-webhooks: building & deploying https://gitlab.wikimedia.org/repos/releng/gitlab-webhooks/-/merge_requests/37 - hitting some build tooling stuff, trying a fix per instructions in the error log
* 17:54 bd808: Unblock 167.56.0.0/13 ([[phab:T422721|T422721]])
* 06:31 hashar: Deleted integration-agent-castor05 Bullseye instance, replaced by integration-agent-castor06 which is on Bookworm # [[phab:T421114|T421114]]
* 06:24 hashar: Deleted integration-agent-qemu-1003 Bullseye image, replaced by integration-agent-qemu-1004 which is on Bookworm # [[phab:T422488|T422488]]
=== 2026-04-07 ===
* 22:25 dduvall: adding new pipelinelib labels to ci nodes ([[phab:T422234|T422234]])
* 20:05 hashar: Triggered a build of https://integration.wikimedia.org/ci/job/mediawiki-core-doxygen/
* 17:06 dduvall: added `Docker` label to `contint` jenkins nodes ([[phab:T422507|T422507]])
* 17:05 dduvall: restored missing `pipelinelib` labels on `integration-agent-docker-` CI hosts ([[phab:T422507|T422507]])
* 16:53 bd808: Unblock 73.0.0.0/8 ([[phab:T422498|T422498]])
* 12:36 hashar: jjb: use $CASTOR_HOST for Quibble success cache. https://gerrit.wikimedia.org/r/1268545 {{!}} This causes the Quibble jobs to use a new instance for the success cache, which is empty # [[phab:T383243|T383243]] [[phab:T421114|T421114]]
* 12:17 hashar: Migrated Castor from integration-castor05 to integration-castor06. Updated CASTOR_HOST in Jenkins and moved the Cinder volume to the new instance # [[phab:T421114|T421114]]
* 11:14 hashar: Added Bookworm based Jenkins agents to the pool Hostnames 1090, 1091, 1092 and 1093 # [[phab:T421114|T421114]]
* 10:09 hashar: Added Bookworm based Jenkins agents to the pool Hostnames 1083 to 1089 # [[phab:T421114|T421114]]
* 07:23 hashar: CI Jenkins: removed `blubber` label from all agents after having moved PipelineLib to use the `Docker` label {{!}} [[phab:T422234|T422234]]
=== 2026-04-06 ===
* 16:01 dancy: Updating docker-pkg files on contint primary for https://gerrit.wikimedia.org/r/c/integration/config/+/1268239
=== 2026-04-03 ===
* 20:17 bd808: Unblock 2.54.0.0/16 ([[phab:T422238|T422238]])
* 17:25 bd808: Unblock 31.18.0.0/16 ([[phab:T422245|T422245]])
* 17:18 bd808: Unblock 2.54.128.0/19 ([[phab:T422238|T422238]])
* 16:18 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1264649 "add Python 3.14 to pywikibot jobs and separate lint tests" {{!}} [[phab:T421723|T421723]]
* 09:26 hashar: integration: nuked pywikibot/core pre-commit cache # [[phab:T422242|T422242]]
* 09:15 hashar: Added Bookworm based Jenkins agents to the pool with label `Docker`. Hostnames are `integration-agent-docker-107*` # [[phab:T421114|T421114]]
* 02:47 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1267398
=== 2026-04-02 ===
* 16:50 thcipriani: restart jenkins
* 15:15 bd808: Unblock 82.216.0.0/16 ([[phab:T421508|T421508]])
* 15:07 bd808: Unblock 95.90.0.0/15 ([[phab:T421485|T421485]])
* 11:19 James_F: Zuul: [oojs/ui] Drop ooui-ruby2.7-rake job, we're abandoning Ruby use there
=== 2026-04-01 ===
* 22:01 bd808: Unblock 109.144.0.0/12 ([[phab:T422019|T422019]])
* 20:16 bd808: Unblock 93.192.0.0/10 ([[phab:T421894|T421894]])
* 19:25 dancy: Updating buildkitd to v0.29.0 in gitlab-cloud-runners (prod) ([[phab:T415284|T415284]])
* 17:57 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/97 ([[phab:T420441|T420441]])
* 17:39 bd808: Unblock 94.134.0.0/15 ([[phab:T421866|T421866]])
* 16:31 dancy: Upgrade buildkit to 0.29.0 in staging gitlab-cloud-runners ([[phab:T415284|T415284]])
* 10:47 taavi: integration-castor05: free up a bit of disk space by deleting cache for AhoCorasick/ CLDRPluralRuleParser/ HtmlFormatter/ RelPath/ RunningStat/ IPSet/
=== 2026-03-30 ===
* 22:01 bd808: Unblock 78.20.0.0/14 ([[phab:T421586|T421586]])
* 21:04 bd808: Unblock 95.88.0.0/15 ([[phab:T421774|T421774]])
* 20:49 bd808: Unblock 95.89.191.0/24 ([[phab:T421774|T421774]])
* 20:29 bd808: Unblock 73.162.0.0/16 ([[phab:T421549|T421549]])
* 13:10 hashar: gerrit: abandon mediawiki/core changes that are 2+years old and are attached to a task (`Bug: Txxxx`)
* 11:37 hashar: Reloaded Zuul to to add 3 persons to the allow list
* 10:43 James_F: Docker: Re-pushing to try to create quibble-coverage 1.16.0-s2
=== 2026-03-27 ===
* 21:00 James_F: Docker: [quibble-bullseye] Drop Python 2 from images
* 11:28 hashar: deployment-prep: removed block for `143.176.0.0/15` and blocked subblock `143.176.0.0/16` instead. This unblocks `143.177.0.0/16` # [[phab:T421420|T421420]]
* 00:18 bd808: Unblock 95.90.238.0/23 ([[phab:T421447|T421447]])
=== 2026-03-26 ===
* 21:25 bd808: Unblock 89.240.0.0/15 ([[phab:T421364|T421364]])
* 21:09 brennen: patchdemo: deploy to production for https://gitlab.wikimedia.org/repos/test-platform/catalyst/patchdemo/-/merge_requests/312
=== 2026-03-25 ===
* 20:41 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256318 [[phab:T421283|T421283]]
* 15:46 dancy: Migrated gitlab-cloud-runners (prod) from nginx-ingress to traefik ([[phab:T420743|T420743]])
* 15:32 dancy: Migrated gitlab-cloud-runners (staging) from nginx-ingress to traefik ([[phab:T420743|T420743]])
* 10:01 hashar: Updating tox Jenkins jobs to add support for Python 3.14 {{!}} https://gerrit.wikimedia.org/r/1260632 {{!}} [[phab:T421209|T421209]]
* 08:40 codders: integration: integration-castor05: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20/
=== 2026-03-24 ===
* 19:40 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1255746
* 15:34 brennen: gitlab1004: manual test run of `configure-projects` with cleared issue allowlist ([[phab:T412882|T412882]])
* 15:26 bd808: Unblock 47.194.0.0/16 ([[phab:T421127|T421127]])
* 12:53 hashar: integration: deleted old Puppet 5 compiler agents from Jenkins ( pcc-worker1014.puppet-diffs.eqiad1.wikimedia.cloud , pcc-worker1015.puppet-diffs.eqiad1.wikimedia.cloud , pcc-worker1016.puppet-diffs.eqiad1.wikimedia.cloud ) # [[phab:T367399|T367399]]
* 07:42 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1259755
=== 2026-03-23 ===
* 15:28 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 90272
=== 2026-03-22 ===
* 14:52 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1258082
* 01:00 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256488
=== 2026-03-21 ===
* 08:10 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256962
* 07:48 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1256946
=== 2026-03-20 ===
* 21:21 bd808: Unblock 103.159.218.0/24 ([[phab:T420530|T420530]])
* 14:59 James_F: Zuul: [mediawiki/extensions/AbuseFilter] Add dependency on CodeMirror, for [[phab:T399673|T399673]]
=== 2026-03-19 ===
* 16:54 Krinkle: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1255777
* 16:01 Krinkle: Hoist l10n-bot rights from labs/tools parent to labs parent to reduce duplication in other labs/ repos
* 15:50 Krinkle: Create labs/xtools repo (branch: main, parent: labs, owner: labs-xtools), ref [[phab:T402086|T402086]]
=== 2026-03-18 ===
* 21:11 dcausse: [[phab:T403775|T403775]]: reindexing all wikis to enable new sorting options
* 21:08 dcausse: restarting opensearch on deployment-cirrussearch(12{{!}}13{{!}}14) instances to pickup new plugin versions
* 14:56 James_F: Zuul: Handle wmf/next the same way as wmf/branch_cut_pretest
* 14:52 James_F: Zuul: [GrowthExperiments] drop duplicate VisualEditor dep
* 14:52 James_F: Zuul: [search/*] Add experimental Java 25 jobs
=== 2026-03-17 ===
* 22:50 James_F: Zuul: [mediawiki/extensions/JsonForms] Add quibble jobs
* 21:27 James_F: Zuul: search: Update opensearch plugins for Java 11/17, for [[phab:T420407|T420407]]
* 20:20 bd808: Resize deployment-sessionstore06 from g4.cores1.ram2.disk20 to g4.cores2.ram4.disk20 ([[phab:T415021|T415021]])
* 16:43 James_F: Zuul: [BlueSpicePermissionManager] Add …ConfigManager & …UserManager deps
* 14:36 James_F: Zuul: [mediawiki/extensions/ArticleGuidance]: Add SpamBlacklist as phan dep, for [[phab:T420015|T420015]]
=== 2026-03-13 ===
* 13:59 andrewbogott: deleting ptr record 117.0.16.172.in-addr.arpa. -- accidental duplicate for deployment-kafka-logging01.deployment-prep.eqiad1.wikimedia.cloud
* 13:04 elukey: re-create kafka-logging-01 in deployment-prep on trixie and Kafka 3.7 (was running on buster)
* 09:13 elukey: upgrade kafka-jumbo and kafka-main to Confluent 7.7 in deployment-prep (pre-requisite before being able to upgrade to Trixie)
=== 2026-03-12 ===
* 21:23 bd808: Hard reboot deployment-sessionstore06 ([[phab:T415021|T415021]])
* 01:14 James_F: Docker: [helm-linter] Bump for Envoy 1.35.9, for [[phab:T419637|T419637]]
=== 2026-03-11 ===
* 16:48 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/MetricsPlatform # [[phab:T417568|T417568]]
* 16:47 James_F: Zuul: [mediawiki/extensions/MetricsPlatform] Archive, for [[phab:T416865|T416865]]
* 11:12 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1250529 "inference-services: Split policy violation CI into separate model jobs." - [[phab:T418832|T418832]]
=== 2026-03-10 ===
* 17:39 dduvall: deployed reggie v1.18.0 to gitlab-cloud-runner production
* 17:11 hashar: Updated MediaWiki coverage jobs so that they now keep "Generate a local configuration by running `composer phpunit:config`" message # [[phab:T419073|T419073]]
* 16:41 dduvall: deployed reggie v1.18.0 to gitlab-cloud-runner staging
* 08:21 codders: integration: integration-castor05: rm -fR /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20
=== 2026-03-09 ===
* 21:53 bd808: Reboot deployment-shellbox01 on the off chance that is makes the new permissions error go away ([[phab:T419440|T419440]])
* 13:13 James_F: Zuul: [mediawiki/extensions/WikiShare] Mark as archived, for [[phab:T413589|T413589]]
* 13:11 James_F: Zuul: [mediawiki/extensions/Memento] Mark as archived, for [[phab:T369991|T369991]]
* 13:10 James_F: Zuul: [mediawiki/extensions/QuickGV] Mark as archived, for [[phab:T413348|T413348]]
* 13:10 James_F: Zuul: [mediawiki/extensions/SemanticImageInput] Mark as archived, for [[phab:T413588|T413588]]
* 13:09 James_F: Zuul: [mediawiki/extensions/SidebarDonateBox] Mark as archived, for [[phab:T413587|T413587]]
* 13:07 James_F: Zuul: [mediawiki/extensions/SemanticSifter] Mark as archived, for [[phab:T413586|T413586]]
* 13:06 James_F: Zuul: [mediawiki/extensions/GoogleAdSense] Mark as archived, for [[phab:T413585|T413585]]
* 13:04 James_F: Zuul: [mediawiki/extensions/SecurityAPI] Mark as archived, for [[phab:T418008|T418008]]
* 12:50 James_F: Zuul: [mediawiki/extensions/CheckUser] Add DiscussionTools dependency
* 12:50 James_F: Zuul: [mediawiki/skins/MinervaNeue] Add dependencies for TestKitchen
* 10:40 hashar: gerrit: mediawiki/vendor: converted `es6` and `es710` branches to tags # [[phab:T417804|T417804]]
* 09:24 hashar: Updating Quibble jobs to 1.16.0 {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1248880 {{!}} [[phab:T417399|T417399]] [[phab:T417409|T417409]] [[phab:T418461|T418461]]
* 09:15 hashar: updating all CI Jenkins jobs using `./jjb-update`
=== 2026-03-06 ===
* 19:46 James_F: Zuul: [mediawiki/services/geoshapes] Mark as archived, for [[phab:T418372|T418372]]
* 16:37 hashar: Building Docker images for Quibble 1.16.0
* 16:31 hashar: Tag Quibble 1.16.0 @ {{Gerrit|0b9db5fe3cabb2cec0b5d44e128bafa917b3b895}} # [[phab:T417399|T417399]] [[phab:T417409|T417409]] [[phab:T418461|T418461]]
* 12:32 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1248411 "jjb, Zuul: vary Wikibase Selenium for release branches" {{!}} [[phab:T418797|T418797]]
* 12:12 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1248409/ "jjb, Zuul: rename wikibase-selenium job for clarity" {{!}} [[phab:T418797|T418797]]
=== 2026-03-05 ===
* 14:41 James_F: Zuul: [mediawiki/skins/MinervaNeue] Add TestKitchen as a dependency for [[phab:T418053|T418053]]
* 08:01 hashar: Reloaded Zuul to rename wikibase-client / wikibase-repo jobs {{!}} https://gerrit.wikimedia.org/r/1238317
* 00:04 James_F: Docker: [quibble-coverage] Use local PHPUnit config, for [[phab:T345481|T345481]]
=== 2026-03-04 ===
* 21:16 James_F: Zuul: [mediawiki/core] Make PHP 8.5 voting on master branch, for [[phab:T411814|T411814]]
* 21:10 James_F: Zuul: [mediawiki/vendor] Make PHP 8.5 voting on master branch, for [[phab:T411814|T411814]]
* 19:48 brennen: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/96 ([[phab:T419004|T419004]])
* 18:50 James_F: Revert "Zuul: [mediawiki/extensions/MobileFrontend] Add ParserMigration dependency", for [[phab:T419043|T419043]]
* 16:23 James_F: Zuul: [mediawiki/services/parsoid] Make PHP 8.4 voting
* 15:37 James_F: Docker: [rake-ruby2.7] Add libffi-dev too, for [[phab:T418463|T418463]]
* 13:59 James_F: Docker: [rake-ruby2.7] Add ruby-ffi for [[phab:T418463|T418463]]
* 13:54 hashar: SIGKILL Zuul cause it can't gracefully stop most probably due to being locked attempting to report back to Gerrit # [[phab:T419009|T419009]]
* 13:49 hashar: Stopping Zuul # [[phab:T419009|T419009]]
* 13:41 hashar: Took a Zuul stack dump on contint1002.wikimedia.org using SIGUSR1 # [[phab:T419009|T419009]]
=== 2026-03-03 ===
* 23:52 James_F: Zuul: [mediawiki/extensions/WikimediaMessages] Drop MetricsPlatform phan dep
* 23:52 James_F: Zuul: [mediawiki/extensions/WikimediaEvents] Drop MetricsPlatform phan dep
=== 2026-03-02 ===
* 22:13 James_F: Zuul: Enforce PHP 8.4 in MW extensions and skins for development branch, for [[phab:T386108|T386108]]
* 14:05 James_F: Zuul: [mediawiki/extensions/MobileFrontend] Add ParserMigration dependency, for [[phab:T415451|T415451]]
* 13:48 James_F: Zuul: […/WikimediaEvents] Drop LoginNotify dependency, now unused, for [[phab:T404334|T404334]]
* 10:16 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/quibble-vendor-mysql-php83-selenium/Cypress/15.8.2/ # [[phab:T418718|T418718]]
=== 2026-02-28 ===
* 21:33 hashar: gerrit: triggering replication to GitHub for all of `mediawiki/skins` # [[phab:T418675|T418675]]
* 21:33 hashar: gerrit: triggering replication to GitHub for all of `mediawiki/extensions` # [[phab:T418675|T418675]]
=== 2026-02-27 ===
* 15:53 dancy: Updating gitlab-cloud-runners (staging and prod) to gitlab-runner 18.9.0.
=== 2026-02-26 ===
* 20:16 James_F: Zuul: Provide a custom, high-priority pipeline just for puppet compiler [[phab:T414621|T414621]]
* 19:32 James_F: Docker: Bump all the PHPs.
* 13:40 hashar: Deployed Jenkins job https://integration.wikimedia.org/ci/job/wikibase-selenium/ # [[phab:T287582|T287582]]
* 00:13 dduvall: forcing replacement of buildkitd helm release in gitlab-cloud-runner prod cluster due to dependency on removed k8s secret ([[phab:T416260|T416260]])
=== 2026-02-25 ===
* 23:50 dduvall: deploying https://gitlab.wikimedia.org/repos/releng/gitlab-cloud-runner/-/merge_requests/552 to gitlab-cloud-runner production cluster ([[phab:T416260|T416260]])
* 14:07 James_F: Zuul: [mediawiki/extensions/CommunityRequests] Add TemplateData dependency, for [[phab:T401638|T401638]]
* 00:08 jeena: no-op testing updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/95
=== 2026-02-24 ===
* 15:55 brennen: devtools: test deploy phab/phorge to test instance ([[phab:T418256|T418256]])
=== 2026-02-23 ===
* 23:07 jeena: Updated development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/92
* 22:43 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/92
* 22:12 bd808: Unblock 191.80.192.0/18 ([[phab:T418132|T418132]])
* 20:26 hashar: Deleted "replication-upstream" Grafana dashboard in favor of a copy/new "replication" one. https://grafana.wikimedia.org/d/RFLS1GsWk/replication-upstream , replaced it by https://grafana.wikimedia.org/d/d4a4da73-c27f-4ce6-a9e5-ab84dd7a4ebb/replication
* 16:29 James_F: Zuul: [3d2png] Add basic Node CI at version 20
=== 2026-02-20 ===
* 21:47 bd808: Unblock 168.184.84.0/24 ([[phab:T418020|T418020]])
* 17:13 bd808: Unblock 122.187.64.0/18 ([[phab:T417964|T417964]])
* 14:35 James_F: Zuul: [mediawiki/extensions/Monstranto] Move out of Wikimedia prod section
=== 2026-02-19 ===
* 18:34 bd808: Unblock 181.98.0.0/16 ([[phab:T417890|T417890]])
* 17:21 James_F: Zuul: [mediawiki/extensions/WikimediaEvents] Add AbuseFilter as a dependency, for [[phab:T417799|T417799]]
* 13:22 hashar: Reloaded Zuul to archive the Cergen repository {{!}} https://gerrit.wikimedia.org/r/c/integration/config/+/1240688 {{!}} [[phab:T417887|T417887]]
=== 2026-02-18 ===
* 20:17 jeena: Updating development images on contint primary for [[phab:T415922|T415922]]
* 19:44 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240360
* 18:40 bd808: Unblock 46.59.0.0/17 ([[phab:T417747|T417747]])
* 17:05 hashar: Regenerating Jenkins jobs with JJB based on https://gerrit.wikimedia.org/r/c/integration/config/+/1240254/
* 17:04 hashar: Added EXT_DEPENDENCIES to Quibble Jenkins jobs parameters so we can manually trigger them from the Web UI using a different set of deps # https://gerrit.wikimedia.org/r/c/integration/config/+/1240254/
* 16:30 hashar: Triggered https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/ with empty Zuul parameters introduced by https://gerrit.wikimedia.org/r/1240333 {{!}} https://integration.wikimedia.org/ci/job/mwcore-phpunit-coverage-master/4893/console
* 15:43 James_F: Zuul: [mediawiki/extensions/ReadingLists] Add EventBus dependency for [[phab:T417706|T417706]]
* 12:15 hashar: zuul-1001.zuul3.eqiad1.wikimedia.cloud: added keepalive=20 to the scheduler Gerrit driver and restarted scheduler container # [[phab:T417497|T417497]]
* 06:58 jeena: Updating development images on contint primary for [[phab:T415922|T415922]]
=== 2026-02-17 ===
* 23:37 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240081
* 23:20 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1240078
* 15:58 brennen: deployed latest phab/phorge wmf/stable to devtools test instance ([[phab:T417657|T417657]])
* 09:01 hashar: Reloaded Zuul to enable php 8.5 testing on utfnormal, php-session-serializer, wikipeg, mediawiki/libs/Dodo, mediawiki/libs/UUID, testing-access-wrapper and translatewiki # [[phab:T406326|T406326]]
=== 2026-02-16 ===
* 15:27 hashar: Manually cleaned some old workspaces on integration-agent-docker-1042
=== 2026-02-12 ===
* 20:07 James_F: Zuul: Enable PHP 8.5 jobs for most MW libraries, for [[phab:T406326|T406326]]
* 19:33 James_F: Docker: [php83] Re-build with upstream's new 8.3.30 release and cascade
* 19:31 James_F: Zuul: Add PHP 8.5 CI job to various things noted as blocked by Phan, for [[phab:T410941|T410941]], [[phab:T406326|T406326]]
* 16:35 Krinkle: Disable publishing noise on tasks from repos Bcp47, clover-diff, ScopedCallback, and IDLeDOM. Ref [[phab:T143162|T143162]]
* 15:53 dancy: Updating development images on contint primary for https://gitlab.wikimedia.org/repos/releng/dev-images/-/merge_requests/87
* 11:21 James_F: Zuul: [mediawiki/libs/shellbox] Add direct Phan job, for [[phab:T416064|T416064]]
=== 2026-02-10 ===
* 20:16 dancy: Rebooted k3s.catalyst-dev (it was unresponsive, but the reboot hasn't helped)
=== 2026-02-09 ===
* 21:58 James_F: Zuul: [mediawiki/tools/phan] Add PHP 8.5 CI job, for [[phab:T410941|T410941]]
* 19:46 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1238006 [[phab:T415680|T415680]]
* 11:51 James_F: Zuul: [mediawiki/extensions/ReadingLists] Drop MetricsPlatform dependency, for [[phab:T414435|T414435]]
=== 2026-02-05 ===
* 17:58 James_F: Zuul: […/WikimediaCustomizations] Add six new dependencies for [[phab:T404334|T404334]]
* 15:35 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1237254
* 15:18 James_F: Zuul: […/OATHAuth] Add dependency and phan dependency on CentralAuth
=== 2026-02-04 ===
* 12:54 James_F: Zuul: [mediawiki/extensions/Petition] Add CLDR dependency
* 10:03 hashar: Restarted Jenkins on releases2003.codfw.wmnet
=== 2026-02-02 ===
* 21:17 hashar: Reloaded Zuul for https://gerrit.wikimedia.org/r/c/integration/config/+/1234926 "re-enable master jobs for some BlueSpice repos - [[phab:T403196|T403196]]"
* 21:05 bd808: Unblock 85.146.0.0/17 ([[phab:T416079|T416079]])
* 19:47 James_F: Zuul: […/WikimediaCustomizations] Add cldr phan dependency, for [[phab:T404334|T404334]]
* 17:33 bd808: Unblock 188.188.0.0/15 ([[phab:T416095|T416095]])
* 17:26 bd808: Unblock 85.94.84.0/22 ([[phab:T416105|T416105]])
* 17:09 bd808: Unblock 94.234.0.0/16 ([[phab:T416165|T416165]])
* 16:51 dancy: Update gitlab-runners to alpine-v18.6.6 ([[phab:T415214|T415214]])
* 16:27 bd808: Unblock 47.231.208.0/21 ([[phab:T416010|T416010]])
* 11:39 James_F: Zuul: […/WikimediaCustomizations] Add five new phan dependencies, for [[phab:T404334|T404334]]
* 09:45 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 58532, 58557
=== 2026-01-31 ===
* 21:49 James_F: Deleted Jenkins's job entry for castor-save-workspace-cache {{Gerrit|6193776}} and this seems to have unstuck things for [[phab:T416078|T416078]]?
* 21:45 James_F: Running `sudo systemctl restart jenkins` on contint for [[phab:T416078|T416078]]
* 21:44 James_F: Fighting [[phab:T416078|T416078]], took integration-castor-5 offline, disconnected, sshed in to kill threads, then reconnected; no change in aspect.
* 19:03 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1235380
=== 2026-01-28 ===
* 21:26 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/WebAuthn # [[phab:T415832|T415832]]
* 21:11 bd808: Unblock 181.160.0.0/15 & 186.40.128.0/17 ([[phab:T415820|T415820]])
* 17:01 bd808: Unblock 102.182.0.0/16 ([[phab:T415782|T415782]])
=== 2026-01-27 ===
* 16:45 James_F: Zuul: Switch skin-quibble template with identical extension-quibble, for [[phab:T402398|T402398]]
* 16:18 James_F: Zuul: [ArticleGuidance] mention it will be in production
* 15:55 James_F: Docker: [quibble-bullseye] Update to Quibble 1.15.0
* 15:12 James_F: Docker: [quibble-coverage] Pass PHPUnit config location explicitly, for [[phab:T395470|T395470]]
* 09:18 hashar: integration: on integration-castor05, deleted caches for old MediaWiki branches
* 09:15 hashar: integration: on pkgbuilder instances, removed Buster cow images, aptcache and hooks. `sudo cumin --force -p 0 'name:pkgbuilder' 'rm -fR /srv/pbuilder/<nowiki>{</nowiki>base-buster-amd64.cow,hooks/buster,aptcache/buster-amd64<nowiki>}</nowiki>'` # [[phab:T397209|T397209]]
* 09:14 hashar: integration: cleaned up old workspaces under /srv/jenkins/workspace
=== 2026-01-26 ===
* 23:27 bd808: Unblock 66.130.0.0/15 ([[phab:T415596|T415596]])
* 22:52 bd808: Unblock 45.16.0.0/12 ([[phab:T415467|T415467]])
* 14:46 hashar: gerrit: changed `operations/software/permissions` project type from `CODE` to `PERMISSIONS` by pointing `HEAD` to `refs/meta/config`
=== 2026-01-22 ===
* 17:36 James_F: Docker: [quibble-coverage] Stop using legacy PHPUnit entrypoint ([[phab:T395470|T395470]]) & Stop excluding Dump/ParserFuzz/Stub groups ([[phab:T415230|T415230]])
* 15:11 James_F: Zuul: [mediawiki/extensions/Math] Add a standalone job, for [[phab:T415230|T415230]]
=== 2026-01-20 ===
* 20:38 bd808: Cherry picked https://gerrit.wikimedia.org/r/c/operations/puppet/+/1229186 ([[phab:T415113|T415113]])
* 19:05 bd808: Rebooted deployment-cache-text08 to see if the mystery haproxy startup failure would go away ([[phab:T415100|T415100]])
* 18:50 bd808: Unblock 152.7.0.0/16 ([[phab:T415100|T415100]])
=== 2026-01-17 ===
* 23:32 ori: beta-scap with `php_l10n: true` completed successfully: https://integration.wikimedia.org/ci/view/Beta/job/beta-scap-sync-world/241466/console. PHP l10n files generated. Reverted local change to scap.cfg.
* 23:26 ori: Temporarily set `php_l10n: true` on deployment-deploy04:/etc/scap.cfg to see if next scap succeeds.
=== 2026-01-16 ===
* 16:33 dancy: Deleting deployment-mx03.deployment-prep ([[phab:T412975|T412975]])
=== 2026-01-15 ===
* 14:50 James_F: jforrester@doc1004:~$ sudo -u doc-uploader rm -rf /srv/doc/cover-extensions/ArticleSummaries/ # [[phab:T413232|T413232]]
=== 2026-01-14 ===
* 17:14 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1226907
* 16:27 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1226893
* 15:57 bd808: Unblock 190.60.63.0/24 ([[phab:T414541|T414541]])
=== 2026-01-13 ===
* 15:04 James_F: Zuul: Make quibble-for-mediawiki-core-vendor-mysql-php84 voting, for [[phab:T386108|T386108]]
=== 2026-01-12 ===
* 21:33 zabe: zabe@deployment-mwmaint03:~$ foreachwiki migrateLinksTable.php --table imagelinks # [[phab:T413668|T413668]]
* 21:06 bd808: Unblock 66.81.168.0/21 ([[phab:T414303|T414303]])
* 17:42 dancy: Turned off instance deployment-prep.deployment-mx03
* 11:44 Lucas_WMDE: ssh integration-castor05.integration.eqiad1.wikimedia.cloud sudo -u jenkins-deploy rm -rf /srv/castor/castor-mw-ext-and-skins/master/mediawiki-node20 # fix failure seen in mediawiki-node20 46331, 46344
=== 2026-01-10 ===
* 21:48 taavi: reload zuul for https://gerrit.wikimedia.org/r/1224782
* 00:25 bd808: Unblock 91.160.0.0/12 ([[phab:T414190|T414190]])
=== 2026-01-09 ===
* 17:33 thcipriani: re-enabling beta update jobs after test bad extension-list [[phab:T411516|T411516]]
* 17:09 thcipriani: disabling beta update jobs to test bad extension-list [[phab:T411516|T411516]])
=== 2026-01-08 ===
* 21:30 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1224815 [[phab:T414136|T414136]]
* 18:24 bd808: Unblock 89.80.0.0/12 ([[phab:T414113|T414113]])
* 15:55 dancy: Upgrading gitlab-runner to v18.5.0 on gitlab-cloud-runners. ([[phab:T414053|T414053]])
=== 2026-01-07 ===
* 23:17 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1082574 https://gerrit.wikimedia.org/r/1224157 https://gerrit.wikimedia.org/r/1224159
* 23:12 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/896311 [[phab:T27482|T27482]]
* 23:06 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1224218
* 17:34 James_F: Zuul: Add new extensions: IssueTrackerLinks, PreviewLinks, and WikiRAG
* 17:34 James_F: Zuul: [labs/tools/heritage] Point to the task to drop 8.1 testing
* 15:09 James_F: Zuul: [labs/tools/heritage] Add testing in PHP 8.2+, not just PHP 8.1
* 15:03 James_F: Zuul: Even for extension-broken, don't offer PHP 8.1 testing
* 15:02 James_F: Zuul: Move quibble experimental sqlite/postgres tests to PHP 8.3
=== 2026-01-06 ===
* 16:57 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1223690 [[phab:T411814|T411814]]
* 16:16 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1223189 [[phab:T411814|T411814]]
* 00:30 bd808: Unblock 85.134.128.0/17 ([[phab:T413755|T413755]])
* 00:02 bd808: Unblock 89.166.128.0/17 ([[phab:T413702|T413702]])
=== 2026-01-05 ===
* 23:57 bd808: Unblock 185.233.104.0/22 ([[phab:T413472|T413472]])
* 23:51 bd808: Unblock 45.62.112.0/21 ([[phab:T413079|T413079]])
* 23:44 bd808: Unblock 85.134.200.0/21 ([[phab:T413067|T413067]])
* 19:03 dancy: Updated buildkitd to v0.26.3 in gitlab-cloud-runners
* 14:27 taavi: reload zuul for {{Gerrit|1223191}}
* 13:57 James_F: Zuul: [mediawiki/php/wmerrors] Enable PHP 8.5 testing, for [[phab:T410921|T410921]]
=== 2026-01-03 ===
* 17:59 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1222709 https://gerrit.wikimedia.org/r/1220388 https://gerrit.wikimedia.org/r/1219140
=== 2026-01-02 ===
* 17:10 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1222597
=== 2026-01-01 ===
* 02:34 Reedy: Reloading Zuul to deploy https://gerrit.wikimedia.org/r/1221644
<noinclude>'''Server Admin Log''' logged from {{IRC|wikimedia-releng}} for [[Nova Resource:Deployment-prep|Beta Cluster]], [[mw:Continuous integration|Continuous integration]] and various other Release Engineering projects.</noinclude>
{{SAL-archives/Release Engineering}}
<noinclude>[[Category:SAL]]</noinclude>
bd7bvckkvrcbc93gq3waa87xxycq9ms
Map of database maintenance
0
449160
2461136
2461095
2026-09-27T00:02:25Z
Dexbot
30554
Bot: Updating the report
2461136
wikitext
text/x-wiki
{{/Header}}
== Today (2026-09-27) ==
== Yesterday (2026-09-26) ==
== Last seven days ==
[[Category:MariaDB]]
iilmwoamxc9rnbqru1dgcetb2sazwc4
Tool:Gitlab-account-approval/Log
116
453906
2461139
2460089
2026-09-27T04:30:37Z
Gitlabaccountapprovalbot
37332
bharu-123 was rejected.
2461139
wikitext
text/x-wiki
<noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude>
=== 2026-09-27 ===
* 04:30 "bharu-123" was rejected (pending since 2026-06-28T04:29:14.159Z).
=== 2026-09-23 ===
* 17:48 "mdh" was rejected (pending since 2026-06-24T17:46:51.726Z).
=== 2026-09-19 ===
* 06:42 "thet1nk3r3r" was rejected (pending since 2026-06-20T06:40:39.066Z).
=== 2026-09-15 ===
* 19:57 "modalya" was rejected (pending since 2026-06-16T19:55:12.189Z).
=== 2026-09-14 ===
* 01:12 [[gitlab:11wb|@11wb]] was approved.
=== 2026-09-13 ===
* 20:54 [[gitlab:mrstarfleetcommand|@mrstarfleetcommand]] was approved.
* 11:36 "nyxira17" was rejected (pending since 2026-06-14T11:35:22.416Z).
=== 2026-09-12 ===
* 19:06 "deuces9ers" was rejected (pending since 2026-06-13T19:04:55.275Z).
=== 2026-09-10 ===
* 08:54 "nehap1706" was rejected (pending since 2026-06-11T08:53:58.819Z).
* 01:48 "tehkittycat" was rejected (pending since 2026-06-11T01:47:45.480Z).
=== 2026-09-09 ===
* 20:12 "lufi2586" was rejected (pending since 2026-06-10T20:09:59.034Z).
=== 2026-09-08 ===
* 02:39 "tantomile" was rejected (pending since 2026-06-09T02:36:22.460Z).
=== 2026-09-07 ===
* 15:45 [[gitlab:dvrted|@dvrted]] was approved.
* 03:00 "jessica22664" was rejected (pending since 2026-06-08T02:58:30.150Z).
=== 2026-09-06 ===
* 21:36 "rp238" was rejected (pending since 2026-06-07T21:33:48.927Z).
* 11:57 "versatiledove231" was rejected (pending since 2026-06-07T11:55:14.557Z).
=== 2026-09-05 ===
* 19:51 "paper9oll" was rejected (pending since 2026-06-06T19:48:46.353Z).
* 04:21 "luminousstaring" was rejected (pending since 2026-06-06T04:20:44.385Z).
=== 2026-09-03 ===
* 08:09 [[gitlab:real68er|@real68er]] was approved.
=== 2026-09-02 ===
* 23:30 [[gitlab:aplucas0703|@aplucas0703]] was approved.
* 16:42 "maximocampo" was rejected (pending since 2026-06-03T16:41:52.538Z).
=== 2026-09-01 ===
* 20:15 "sinaider2000" was rejected (pending since 2026-06-02T20:12:40.894Z).
* 19:03 "t-rt15" was rejected (pending since 2026-06-02T19:02:56.111Z).
* 04:21 "naomi723" was rejected (pending since 2026-06-02T04:20:21.044Z).
=== 2026-08-31 ===
* 20:06 "mccio" was rejected (pending since 2026-06-01T20:05:09.900Z).
* 10:51 [[gitlab:hgzh3|@hgzh3]] was approved.
=== 2026-08-27 ===
* 02:18 [[gitlab:madmax|@madmax]] was approved.
=== 2026-08-26 ===
* 21:42 [[gitlab:mmns21|@mmns21]] was approved.
=== 2026-08-25 ===
* 17:27 [[gitlab:skim|@skim]] was approved.
* 14:57 [[gitlab:wilfredor2|@wilfredor2]] was approved.
* 08:03 [[gitlab:noahmvf|@noahmvf]] was approved.
=== 2026-08-23 ===
* 05:30 "amidewicki" was rejected (pending since 2026-05-24T05:29:29.216Z).
* 01:54 "westcodragen07" was rejected (pending since 2026-05-24T01:53:53.396Z).
=== 2026-08-22 ===
* 15:21 [[gitlab:laraibhasan|@laraibhasan]] was approved.
=== 2026-08-20 ===
* 17:18 [[gitlab:jatwa|@jatwa]] was approved.
=== 2026-08-18 ===
* 11:33 "wladek92" was rejected (pending since 2026-05-19T11:30:32.679Z).
=== 2026-08-17 ===
* 18:39 [[gitlab:erlanger16e|@erlanger16e]] was approved.
* 18:33 "ashishkaranam" was rejected (pending since 2026-05-18T18:32:18.813Z).
* 12:48 [[gitlab:lfs|@lfs]] was approved.
=== 2026-08-16 ===
* 15:42 [[gitlab:pcendrer|@pcendrer]] was approved.
=== 2026-08-15 ===
* 17:33 [[gitlab:flammablepizza|@flammablepizza]] was approved.
=== 2026-08-14 ===
* 13:00 [[gitlab:pacmand|@pacmand]] was approved.
=== 2026-08-12 ===
* 13:21 [[gitlab:doublechek|@doublechek]] was approved.
=== 2026-08-11 ===
* 14:15 [[gitlab:mmilenkovicwmf|@mmilenkovicwmf]] was approved.
=== 2026-08-10 ===
* 17:18 [[gitlab:mustafe|@mustafe]] was approved.
* 10:33 "c8263a20" was rejected (pending since 2026-05-11T10:32:40.353Z).
=== 2026-08-09 ===
* 21:18 [[gitlab:iamnetx|@iamnetx]] was approved.
* 20:09 "yirba" was rejected (pending since 2026-05-10T20:07:07.738Z).
* 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z).
=== 2026-08-08 ===
* 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved.
=== 2026-08-07 ===
* 18:21 [[gitlab:gturkington|@gturkington]] was approved.
* 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z).
=== 2026-07-30 ===
* 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z).
=== 2026-07-29 ===
* 10:03 [[gitlab:piastu|@piastu]] was approved.
* 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z).
=== 2026-07-28 ===
* 16:51 [[gitlab:for-each-next|@for-each-next]] was approved.
* 09:21 [[gitlab:gka|@gka]] was approved.
=== 2026-07-27 ===
* 08:12 [[gitlab:cambob|@cambob]] was approved.
=== 2026-07-26 ===
* 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z).
=== 2026-07-24 ===
* 10:54 [[gitlab:ysogo|@ysogo]] was approved.
* 09:39 [[gitlab:pankaj199|@pankaj199]] was approved.
=== 2026-07-23 ===
* 12:30 [[gitlab:fermiboson|@fermiboson]] was approved.
* 09:21 [[gitlab:slashme|@slashme]] was approved.
=== 2026-07-22 ===
* 15:54 [[gitlab:lmedley|@lmedley]] was approved.
* 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z).
* 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved.
* 08:30 [[gitlab:nabbegat|@nabbegat]] was approved.
* 08:30 [[gitlab:plyd|@plyd]] was approved.
* 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z).
=== 2026-07-21 ===
* 15:51 [[gitlab:panieravide|@panieravide]] was approved.
* 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved.
* 13:48 [[gitlab:yru|@yru]] was approved.
* 13:09 [[gitlab:deevad|@deevad]] was approved.
* 12:54 [[gitlab:ctdo17|@ctdo17]] was approved.
* 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved.
* 12:21 [[gitlab:nadiantara|@nadiantara]] was approved.
* 12:18 [[gitlab:wijltcher|@wijltcher]] was approved.
* 10:33 [[gitlab:nivopol|@nivopol]] was approved.
* 10:30 [[gitlab:johlig|@johlig]] was approved.
* 10:30 [[gitlab:majicita|@majicita]] was approved.
* 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved.
* 10:09 [[gitlab:francyskus|@francyskus]] was approved.
* 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved.
=== 2026-07-20 ===
* 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z).
* 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved.
* 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z).
=== 2026-07-19 ===
* 12:21 [[gitlab:boivie|@boivie]] was approved.
=== 2026-07-18 ===
* 16:09 [[gitlab:pharos|@pharos]] was approved.
* 15:45 [[gitlab:priyankar22|@priyankar22]] was approved.
* 15:30 [[gitlab:sisyph|@sisyph]] was approved.
=== 2026-07-13 ===
* 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved.
=== 2026-07-12 ===
* 09:27 [[gitlab:smk|@smk]] was approved.
=== 2026-07-11 ===
* 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z).
* 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z).
=== 2026-07-10 ===
* 14:42 [[gitlab:akaza24|@akaza24]] was approved.
* 12:57 [[gitlab:kormisk|@kormisk]] was approved.
=== 2026-07-07 ===
* 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved.
* 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z).
=== 2026-07-06 ===
* 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z).
=== 2026-07-02 ===
* 15:48 [[gitlab:tekneos|@tekneos]] was approved.
=== 2026-07-01 ===
* 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved.
* 14:15 [[gitlab:vadymts1|@vadymts1]] was approved.
* 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z).
=== 2026-06-30 ===
* 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z).
=== 2026-06-29 ===
* 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved.
=== 2026-06-28 ===
* 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z).
* 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z).
* 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z).
* 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z).
=== 2026-06-26 ===
* 09:39 [[gitlab:lakshita28|@lakshita28]] was approved.
* 07:30 [[gitlab:reeti|@reeti]] was approved.
* 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved.
* 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z).
* 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved.
* 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved.
=== 2026-06-25 ===
* 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved.
* 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z).
=== 2026-06-24 ===
* 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved.
* 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z).
=== 2026-06-22 ===
* 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved.
* 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved.
* 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z).
=== 2026-06-21 ===
* 13:18 [[gitlab:egonw|@egonw]] was approved.
=== 2026-06-20 ===
* 10:21 [[gitlab:tways2017|@tways2017]] was approved.
=== 2026-06-19 ===
* 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z).
* 04:12 [[gitlab:claudio|@claudio]] was approved.
=== 2026-06-18 ===
* 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z).
* 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved.
=== 2026-06-17 ===
* 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z).
* 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z).
=== 2026-06-15 ===
* 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved.
* 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z).
* 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z).
* 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z).
=== 2026-06-14 ===
* 17:45 [[gitlab:surajseth520|@surajseth520]] was approved.
* 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z).
=== 2026-06-11 ===
* 11:48 [[gitlab:cadddr|@cadddr]] was approved.
* 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z).
* 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved.
=== 2026-06-10 ===
* 07:03 [[gitlab:dmiranda|@dmiranda]] was approved.
=== 2026-06-09 ===
* 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved.
* 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved.
* 12:51 [[gitlab:ekrem|@ekrem]] was approved.
=== 2026-06-08 ===
* 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z).
* 16:15 [[gitlab:ahonc|@ahonc]] was approved.
* 12:57 [[gitlab:rainmonger|@rainmonger]] was approved.
=== 2026-06-07 ===
* 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z).
* 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z).
* 02:39 [[gitlab:launchpad|@launchpad]] was approved.
=== 2026-06-06 ===
* 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z).
* 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z).
=== 2026-06-04 ===
* 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z).
=== 2026-06-03 ===
* 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved.
=== 2026-06-02 ===
* 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z).
* 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z).
* 05:48 [[gitlab:tmoney|@tmoney]] was approved.
=== 2026-06-01 ===
* 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z).
* 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z).
=== 2026-05-29 ===
* 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z).
=== 2026-05-27 ===
* 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z).
* 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z).
* 14:33 [[gitlab:thilio|@thilio]] was approved.
=== 2026-05-26 ===
* 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z).
=== 2026-05-25 ===
* 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z).
* 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z).
* 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z).
=== 2026-05-24 ===
* 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z).
=== 2026-05-23 ===
* 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved.
* 04:21 [[gitlab:staraction|@staraction]] was approved.
=== 2026-05-22 ===
* 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z).
* 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z).
=== 2026-05-21 ===
* 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z).
* 16:27 [[gitlab:renovatebot|@renovatebot]] was approved.
* 16:06 [[gitlab:gkm563|@gkm563]] was approved.
=== 2026-05-20 ===
* 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z).
=== 2026-05-18 ===
* 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z).
* 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved.
=== 2026-05-14 ===
* 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved.
=== 2026-05-13 ===
* 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z).
=== 2026-05-12 ===
* 18:06 [[gitlab:psubhashish|@psubhashish]] was approved.
* 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z).
* 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z).
=== 2026-05-11 ===
* 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z).
=== 2026-05-10 ===
* 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z).
* 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z).
=== 2026-05-09 ===
* 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z).
=== 2026-05-08 ===
* 09:24 [[gitlab:cwilliams|@cwilliams]] was approved.
=== 2026-05-07 ===
* 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z).
=== 2026-05-06 ===
* 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z).
* 08:09 [[gitlab:neriah|@neriah]] was approved.
* 06:27 [[gitlab:status401|@status401]] was approved.
=== 2026-05-03 ===
* 09:54 [[gitlab:anilk|@anilk]] was approved.
=== 2026-05-02 ===
* 17:54 [[gitlab:sweil|@sweil]] was approved.
* 17:00 [[gitlab:aoppo|@aoppo]] was approved.
=== 2026-05-01 ===
* 21:18 [[gitlab:dawalda|@dawalda]] was approved.
=== 2026-04-30 ===
* 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z).
* 20:54 [[gitlab:tfmorris|@tfmorris]] was approved.
* 17:33 [[gitlab:uyen|@uyen]] was approved.
* 07:39 [[gitlab:mahveotm|@mahveotm]] was approved.
* 06:36 [[gitlab:leo321|@leo321]] was approved.
=== 2026-04-29 ===
* 02:27 [[gitlab:dw31415|@dw31415]] was approved.
=== 2026-04-28 ===
* 23:09 [[gitlab:dtorsani|@dtorsani]] was approved.
=== 2026-04-27 ===
* 23:42 [[gitlab:quinlan|@quinlan]] was approved.
* 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved.
=== 2026-04-26 ===
* 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z).
* 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z).
=== 2026-04-25 ===
* 20:24 [[gitlab:maldaxura|@maldaxura]] was approved.
* 14:33 [[gitlab:sirtobi|@sirtobi]] was approved.
* 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z).
=== 2026-04-24 ===
* 22:06 [[gitlab:arcstur|@arcstur]] was approved.
=== 2026-04-22 ===
* 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z).
* 22:18 [[gitlab:egezort|@egezort]] was approved.
* 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z).
=== 2026-04-20 ===
* 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z).
=== 2026-04-19 ===
* 02:54 [[gitlab:neoact|@neoact]] was approved.
=== 2026-04-18 ===
* 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved.
=== 2026-04-17 ===
* 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z).
=== 2026-04-15 ===
* 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z).
=== 2026-04-14 ===
* 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z).
* 04:51 [[gitlab:canley|@canley]] was approved.
* 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z).
=== 2026-04-13 ===
* 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved.
=== 2026-04-11 ===
* 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z).
=== 2026-04-10 ===
* 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z).
=== 2026-04-07 ===
* 17:03 [[gitlab:supnabla|@supnabla]] was approved.
=== 2026-04-06 ===
* 20:00 [[gitlab:laerdon|@laerdon]] was approved.
* 19:21 [[gitlab:ljq3|@ljq3]] was approved.
=== 2026-04-04 ===
* 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z).
=== 2026-04-02 ===
* 05:30 [[gitlab:mbh1|@mbh1]] was approved.
=== 2026-04-01 ===
* 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z).
* 12:12 [[gitlab:amorii0|@amorii0]] was approved.
=== 2026-03-31 ===
* 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z).
=== 2026-03-30 ===
* 15:36 [[gitlab:atsuko|@atsuko]] was approved.
=== 2026-03-29 ===
* 11:36 [[gitlab:giftcup|@giftcup]] was approved.
=== 2026-03-28 ===
* 14:51 [[gitlab:janeeva1|@janeeva1]] was approved.
=== 2026-03-26 ===
* 13:36 [[gitlab:saiphani02|@saiphani02]] was approved.
* 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved.
=== 2026-03-25 ===
* 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z).
* 02:18 [[gitlab:viztor|@viztor]] was approved.
=== 2026-03-24 ===
* 23:18 [[gitlab:maryyann|@maryyann]] was approved.
* 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved.
* 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved.
* 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z).
=== 2026-03-23 ===
* 14:51 [[gitlab:komla|@komla]] was approved.
* 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z).
* 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z).
=== 2026-03-20 ===
* 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z).
* 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z).
* 09:57 [[gitlab:nethahussain|@nethahussain]] was approved.
* 09:27 [[gitlab:piewriter|@piewriter]] was approved.
* 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved.
=== 2026-03-19 ===
* 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z).
=== 2026-03-18 ===
* 20:15 [[gitlab:martinmystere|@martinmystere]] was approved.
=== 2026-03-17 ===
* 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z).
=== 2026-03-16 ===
* 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z).
* 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z).
=== 2026-03-14 ===
* 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z).
* 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z).
* 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved.
=== 2026-03-12 ===
* 20:24 [[gitlab:11wb|@11wb]] was approved.
* 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved.
=== 2026-03-10 ===
* 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved.
=== 2026-03-06 ===
* 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z).
=== 2026-03-04 ===
* 20:54 [[gitlab:elphie|@elphie]] was approved.
* 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z).
* 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z).
=== 2026-03-02 ===
* 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z).
* 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved.
=== 2026-03-01 ===
* 10:51 [[gitlab:clintacc|@clintacc]] was approved.
=== 2026-02-28 ===
* 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z).
* 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z).
=== 2026-02-27 ===
* 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z).
=== 2026-02-26 ===
* 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z).
* 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z).
=== 2026-02-25 ===
* 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z).
=== 2026-02-24 ===
* 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z).
=== 2026-02-23 ===
* 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z).
* 12:33 [[gitlab:robertsky|@robertsky]] was approved.
=== 2026-02-22 ===
* 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z).
=== 2026-02-19 ===
* 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z).
=== 2026-02-18 ===
* 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z).
* 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z).
=== 2026-02-17 ===
* 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z).
=== 2026-02-16 ===
* 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z).
=== 2026-02-15 ===
* 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z).
=== 2026-02-14 ===
* 15:15 [[gitlab:mess|@mess]] was approved.
=== 2026-02-13 ===
* 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z).
* 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved.
=== 2026-02-12 ===
* 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z).
* 03:33 [[gitlab:anyehansen|@anyehansen]] was approved.
* 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved.
=== 2026-02-10 ===
* 13:18 [[gitlab:db111|@db111]] was approved.
=== 2026-02-09 ===
* 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z).
=== 2026-02-06 ===
* 20:54 [[gitlab:gillux|@gillux]] was approved.
* 09:09 [[gitlab:lih|@lih]] was approved.
=== 2026-01-31 ===
* 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved.
=== 2026-01-28 ===
* 14:30 [[gitlab:ademola|@ademola]] was approved.
* 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z).
=== 2026-01-26 ===
* 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z).
=== 2026-01-25 ===
* 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z).
=== 2026-01-24 ===
* 21:15 [[gitlab:wiegels|@wiegels]] was approved.
* 06:30 [[gitlab:blaquans|@blaquans]] was approved.
=== 2026-01-23 ===
* 16:27 [[gitlab:lerickson|@lerickson]] was approved.
* 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z).
=== 2026-01-22 ===
* 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z).
=== 2026-01-21 ===
* 17:30 [[gitlab:otcenas11|@otcenas11]] was approved.
=== 2026-01-19 ===
* 21:48 [[gitlab:amdrel|@amdrel]] was approved.
* 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z).
=== 2026-01-18 ===
* 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z).
* 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z).
=== 2026-01-16 ===
* 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z).
* 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z).
=== 2026-01-15 ===
* 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z).
* 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z).
* 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z).
=== 2026-01-13 ===
* 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved.
=== 2026-01-12 ===
* 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z).
=== 2026-01-10 ===
* 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z).
=== 2026-01-09 ===
* 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z).
=== 2026-01-08 ===
* 19:42 [[gitlab:tbodt|@tbodt]] was approved.
* 13:57 [[gitlab:martynranyard|@martynranyard]] was approved.
=== 2026-01-07 ===
* 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved.
* 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z).
* 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z).
* 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z).
* 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z).
* 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z).
* 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z).
* 01:15 [[gitlab:tungstalite|@tungstalite]] was approved.
=== 2026-01-06 ===
* 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z).
* 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z).
* 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z).
=== 2026-01-04 ===
* 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved.
* 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z).
* 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z).
=== 2026-01-03 ===
* 20:42 [[gitlab:apehitkey|@apehitkey]] was approved.
* 18:00 [[gitlab:jeremyb|@jeremyb]] was approved.
* 14:09 [[gitlab:twelephant|@twelephant]] was approved.
=== 2026-01-01 ===
* 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z).
=== 2025-12-30 ===
* 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z).
=== 2025-12-29 ===
* 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z).
* 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z).
=== 2025-12-26 ===
* 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z).
=== 2025-12-25 ===
* 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z).
=== 2025-12-23 ===
* 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z).
=== 2025-12-22 ===
* 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z).
=== 2025-12-20 ===
* 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z).
=== 2025-12-19 ===
* 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z).
* 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z).
=== 2025-12-18 ===
* 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved.
=== 2025-12-17 ===
* 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved.
* 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z).
* 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z).
=== 2025-12-16 ===
* 09:39 [[gitlab:melcatherine|@melcatherine]] was approved.
* 08:54 [[gitlab:leila237|@leila237]] was approved.
=== 2025-12-15 ===
* 18:27 [[gitlab:royalsailor|@royalsailor]] was approved.
* 09:39 [[gitlab:olaf8940|@olaf8940]] was approved.
* 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z).
=== 2025-12-14 ===
* 20:21 [[gitlab:essa237|@essa237]] was approved.
* 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved.
=== 2025-12-13 ===
* 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z).
* 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z).
=== 2025-12-12 ===
* 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z).
* 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z).
=== 2025-12-10 ===
* 17:30 [[gitlab:itsmoon|@itsmoon]] was approved.
=== 2025-12-09 ===
* 15:42 [[gitlab:mercy-o|@mercy-o]] was approved.
=== 2025-12-06 ===
* 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z).
* 11:27 [[gitlab:ikhitron|@ikhitron]] was approved.
=== 2025-12-01 ===
* 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z).
=== 2025-11-30 ===
* 21:06 [[gitlab:habs|@habs]] was approved.
=== 2025-11-29 ===
* 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z).
* 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved.
=== 2025-11-24 ===
* 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z).
* 02:24 [[gitlab:yaron|@yaron]] was approved.
=== 2025-11-20 ===
* 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z).
=== 2025-11-17 ===
* 21:09 [[gitlab:ankita97531|@ankita97531]] was approved.
=== 2025-11-16 ===
* 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z).
* 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z).
=== 2025-11-15 ===
* 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z).
=== 2025-11-13 ===
* 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z).
=== 2025-11-11 ===
* 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved.
=== 2025-11-10 ===
* 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved.
=== 2025-11-09 ===
* 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z).
=== 2025-11-07 ===
* 22:00 [[gitlab:forzagreen|@forzagreen]] was approved.
=== 2025-11-06 ===
* 16:57 [[gitlab:rsilvola|@rsilvola]] was approved.
=== 2025-11-04 ===
* 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved.
=== 2025-11-03 ===
* 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z).
=== 2025-11-01 ===
* 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z).
=== 2025-10-31 ===
* 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved.
* 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z).
=== 2025-10-30 ===
* 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved.
=== 2025-10-29 ===
* 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z).
=== 2025-10-28 ===
* 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved.
=== 2025-10-27 ===
* 07:51 [[gitlab:jmoore111|@jmoore111]] was approved.
=== 2025-10-25 ===
* 21:09 [[gitlab:valor|@valor]] was approved.
* 21:03 [[gitlab:booksmurf|@booksmurf]] was approved.
* 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z).
=== 2025-10-24 ===
* 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z).
=== 2025-10-22 ===
* 20:54 [[gitlab:janewanga|@janewanga]] was approved.
* 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z).
* 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z).
=== 2025-10-21 ===
* 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z).
* 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved.
=== 2025-10-17 ===
* 13:21 [[gitlab:blegodwin|@blegodwin]] was approved.
=== 2025-10-16 ===
* 14:51 [[gitlab:bahago|@bahago]] was approved.
* 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z).
* 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z).
=== 2025-10-15 ===
* 13:48 [[gitlab:adwivedii|@adwivedii]] was approved.
* 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved.
* 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z).
* 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z).
=== 2025-10-14 ===
* 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved.
* 15:33 [[gitlab:sadrettin|@sadrettin]] was approved.
* 14:18 [[gitlab:tmwyk|@tmwyk]] was approved.
* 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z).
=== 2025-10-13 ===
* 16:09 [[gitlab:atlas0007|@atlas0007]] was approved.
=== 2025-10-11 ===
* 17:42 [[gitlab:techwizzie|@techwizzie]] was approved.
=== 2025-10-10 ===
* 19:03 [[gitlab:miiswom|@miiswom]] was approved.
* 16:06 [[gitlab:ninatakang|@ninatakang]] was approved.
=== 2025-10-09 ===
* 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved.
* 14:21 [[gitlab:lebogang|@lebogang]] was approved.
* 14:15 [[gitlab:kimondorose|@kimondorose]] was approved.
* 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved.
* 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved.
* 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved.
* 13:45 [[gitlab:system625|@system625]] was approved.
* 13:45 [[gitlab:rolalove|@rolalove]] was approved.
* 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved.
* 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved.
* 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved.
* 13:36 [[gitlab:aradhya|@aradhya]] was approved.
* 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z).
=== 2025-10-08 ===
* 23:36 [[gitlab:sopzy|@sopzy]] was approved.
* 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved.
* 19:39 [[gitlab:levon003|@levon003]] was approved.
* 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved.
* 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved.
* 13:36 [[gitlab:chumzine|@chumzine]] was approved.
* 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved.
* 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved.
* 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z).
* 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z).
=== 2025-10-07 ===
* 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved.
* 21:30 [[gitlab:rona|@rona]] was approved.
* 21:09 [[gitlab:sandijigs|@sandijigs]] was approved.
* 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z).
* 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z).
=== 2025-10-06 ===
* 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z).
=== 2025-10-05 ===
* 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z).
=== 2025-10-04 ===
* 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z).
=== 2025-10-03 ===
* 06:12 [[gitlab:javiermonton|@javiermonton]] was approved.
=== 2025-10-02 ===
* 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z).
=== 2025-10-01 ===
* 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z).
* 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z).
=== 2025-09-30 ===
* 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved.
* 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z).
=== 2025-09-28 ===
* 08:54 [[gitlab:pexerik|@pexerik]] was approved.
=== 2025-09-27 ===
* 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved.
=== 2025-09-26 ===
* 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z).
* 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved.
* 13:12 [[gitlab:spushpit|@spushpit]] was approved.
=== 2025-09-20 ===
* 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z).
=== 2025-09-16 ===
* 05:39 [[gitlab:deepchirp|@deepchirp]] was approved.
=== 2025-09-15 ===
* 22:00 [[gitlab:noisk8|@noisk8]] was approved.
* 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z).
=== 2025-09-13 ===
* 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z).
* 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved.
* 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved.
=== 2025-09-12 ===
* 17:12 [[gitlab:suyash23|@suyash23]] was approved.
* 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z).
=== 2025-09-10 ===
* 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z).
=== 2025-09-09 ===
* 17:03 [[gitlab:vwf|@vwf]] was approved.
* 06:36 [[gitlab:cactusisme|@cactusisme]] was approved.
=== 2025-09-08 ===
* 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z).
* 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z).
* 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z).
=== 2025-09-06 ===
* 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved.
=== 2025-09-05 ===
* 21:45 [[gitlab:sakshi2|@sakshi2]] was approved.
* 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z).
* 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z).
=== 2025-09-04 ===
* 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z).
* 19:00 [[gitlab:perry|@perry]] was approved.
* 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z).
=== 2025-09-02 ===
* 05:48 [[gitlab:aliu|@aliu]] was approved.
=== 2025-08-29 ===
* 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z).
=== 2025-08-28 ===
* 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z).
=== 2025-08-26 ===
* 19:03 [[gitlab:dikkulah|@dikkulah]] was approved.
=== 2025-08-22 ===
* 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved.
=== 2025-08-21 ===
* 07:39 [[gitlab:yuka|@yuka]] was approved.
=== 2025-08-19 ===
* 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved.
=== 2025-08-17 ===
* 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z).
=== 2025-08-15 ===
* 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z).
=== 2025-08-11 ===
* 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z).
=== 2025-08-10 ===
* 13:15 [[gitlab:dactyl|@dactyl]] was approved.
=== 2025-08-09 ===
* 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z).
=== 2025-08-08 ===
* 14:33 [[gitlab:josefanthony|@josefanthony]] was approved.
=== 2025-08-07 ===
* 23:42 [[gitlab:robins7|@robins7]] was approved.
* 21:42 [[gitlab:pols12|@pols12]] was approved.
* 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z).
* 14:57 [[gitlab:alvindulle|@alvindulle]] was approved.
* 14:45 [[gitlab:xentos|@xentos]] was approved.
* 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z).
* 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z).
=== 2025-08-06 ===
* 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z).
* 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z).
=== 2025-08-05 ===
* 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z).
=== 2025-08-01 ===
* 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z).
=== 2025-07-31 ===
* 02:30 [[gitlab:ads|@ads]] was approved.
=== 2025-07-27 ===
* 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z).
* 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved.
* 10:17 [[gitlab:fuzzew|@fuzzew]] was approved.
* 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved.
* 05:48 [[gitlab:ecoholic|@ecoholic]] was approved.
=== 2025-07-26 ===
* 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved.
* 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved.
* 11:48 [[gitlab:hridyakk|@hridyakk]] was approved.
* 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved.
* 11:45 [[gitlab:binetaa|@binetaa]] was approved.
* 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved.
* 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved.
* 10:21 [[gitlab:flykrth|@flykrth]] was approved.
* 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved.
* 10:21 [[gitlab:aadhi13|@aadhi13]] was approved.
* 10:21 [[gitlab:malavikam05|@malavikam05]] was approved.
* 10:18 [[gitlab:nf609|@nf609]] was approved.
* 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved.
* 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved.
=== 2025-07-25 ===
* 09:49 [[gitlab:kasyap9|@kasyap9]] was approved.
* 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved.
=== 2025-07-24 ===
* 19:36 [[gitlab:madutgn|@madutgn]] was approved.
=== 2025-07-23 ===
* 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved.
=== 2025-07-22 ===
* 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved.
=== 2025-07-21 ===
* 17:30 [[gitlab:asadiqui|@asadiqui]] was approved.
* 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved.
* 04:27 [[gitlab:damian|@damian]] was approved.
=== 2025-07-20 ===
* 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z).
=== 2025-07-17 ===
* 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved.
* 13:45 [[gitlab:envlh|@envlh]] was approved.
=== 2025-07-14 ===
* 10:24 [[gitlab:missguru|@missguru]] was approved.
* 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z).
=== 2025-07-13 ===
* 01:01 [[gitlab:l235|@l235]] was approved.
=== 2025-07-11 ===
* 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z).
=== 2025-07-06 ===
* 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z).
=== 2025-07-05 ===
* 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z).
* 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z).
=== 2025-07-04 ===
* 03:15 [[gitlab:galstar|@galstar]] was approved.
=== 2025-07-02 ===
* 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z).
=== 2025-06-29 ===
* 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z).
=== 2025-06-28 ===
* 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z).
=== 2025-06-27 ===
* 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z).
* 22:21 [[gitlab:rluts|@rluts]] was approved.
=== 2025-06-26 ===
* 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z).
=== 2025-06-24 ===
* 17:42 [[gitlab:devjadiya|@devjadiya]] was approved.
* 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z).
=== 2025-06-21 ===
* 00:48 [[gitlab:vriaa|@vriaa]] was approved.
=== 2025-06-18 ===
* 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z).
=== 2025-06-17 ===
* 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z).
* 00:27 [[gitlab:eggroll97|@eggroll97]] was approved.
=== 2025-06-14 ===
* 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z).
=== 2025-06-13 ===
* 16:09 [[gitlab:supergrey|@supergrey]] was approved.
* 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z).
* 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved.
* 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z).
=== 2025-06-12 ===
* 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z).
=== 2025-06-09 ===
* 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z).
* 19:27 [[gitlab:reggycelly|@reggycelly]] was approved.
* 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z).
* 13:21 [[gitlab:greenreaper|@greenreaper]] was approved.
* 09:33 [[gitlab:mmta|@mmta]] was approved.
* 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z).
=== 2025-06-08 ===
* 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z).
=== 2025-06-06 ===
* 11:06 [[gitlab:olea|@olea]] was approved.
=== 2025-06-05 ===
* 20:33 [[gitlab:encodedwp|@encodedwp]] was approved.
* 15:00 [[gitlab:toluayo|@toluayo]] was approved.
* 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved.
* 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z).
=== 2025-06-03 ===
* 21:27 [[gitlab:wewakey|@wewakey]] was approved.
* 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z).
* 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z).
=== 2025-06-02 ===
* 12:01 [[gitlab:jaimedes|@jaimedes]] was approved.
=== 2025-05-30 ===
* 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z).
* 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved.
* 10:06 [[gitlab:gpsleo|@gpsleo]] was approved.
=== 2025-05-29 ===
* 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved.
=== 2025-05-28 ===
* 02:57 [[gitlab:saper|@saper]] was approved.
=== 2025-05-27 ===
* 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved.
* 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z).
=== 2025-05-26 ===
* 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z).
=== 2025-05-21 ===
* 11:12 [[gitlab:guilherme|@guilherme]] was approved.
=== 2025-05-19 ===
* 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved.
=== 2025-05-18 ===
* 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z).
=== 2025-05-17 ===
* 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z).
=== 2025-05-16 ===
* 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved.
=== 2025-05-15 ===
* 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z).
=== 2025-05-14 ===
* 17:42 [[gitlab:llugo|@llugo]] was approved.
=== 2025-05-13 ===
* 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z).
=== 2025-05-11 ===
* 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z).
* 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z).
* 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z).
=== 2025-05-09 ===
* 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved.
=== 2025-05-08 ===
* 08:18 [[gitlab:mengeditch|@mengeditch]] was approved.
=== 2025-05-07 ===
* 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z).
=== 2025-05-06 ===
* 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z).
* 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved.
=== 2025-05-02 ===
* 12:24 [[gitlab:tohaomg|@tohaomg]] was approved.
* 11:48 [[gitlab:mavrikant|@mavrikant]] was approved.
* 11:45 [[gitlab:daanvr|@daanvr]] was approved.
=== 2025-05-01 ===
* 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z).
=== 2025-04-30 ===
* 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z).
=== 2025-04-29 ===
* 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z).
=== 2025-04-26 ===
* 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z).
=== 2025-04-25 ===
* 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z).
* 15:15 [[gitlab:yiming|@yiming]] was approved.
* 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z).
=== 2025-04-23 ===
* 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z).
* 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z).
=== 2025-04-22 ===
* 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z).
* 14:30 [[gitlab:kaimamin|@kaimamin]] was approved.
* 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z).
=== 2025-04-21 ===
* 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z).
=== 2025-04-18 ===
* 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved.
=== 2025-04-16 ===
* 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z).
=== 2025-04-15 ===
* 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z).
=== 2025-04-14 ===
* 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved.
=== 2025-04-11 ===
* 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved.
=== 2025-04-10 ===
* 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z).
=== 2025-04-09 ===
* 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z).
=== 2025-04-06 ===
* 15:36 [[gitlab:elph|@elph]] was approved.
=== 2025-04-02 ===
* 10:33 [[gitlab:ozge|@ozge]] was approved.
=== 2025-03-31 ===
* 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z).
* 15:18 [[gitlab:danyya|@danyya]] was approved.
=== 2025-03-28 ===
* 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved.
* 15:54 [[gitlab:ilanen1|@ilanen1]] was approved.
=== 2025-03-25 ===
* 19:27 [[gitlab:irfo|@irfo]] was approved.
* 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved.
* 04:33 [[gitlab:paul26|@paul26]] was approved.
* 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z).
=== 2025-03-24 ===
* 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z).
=== 2025-03-23 ===
* 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z).
=== 2025-03-22 ===
* 09:45 [[gitlab:fjmustak|@fjmustak]] was approved.
=== 2025-03-20 ===
* 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z).
* 17:03 [[gitlab:alien4444|@alien4444]] was approved.
* 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved.
=== 2025-03-19 ===
* 22:57 [[gitlab:r1f4t|@r1f4t]] was approved.
* 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z).
* 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved.
=== 2025-03-18 ===
* 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z).
* 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z).
=== 2025-03-15 ===
* 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z).
* 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z).
=== 2025-03-14 ===
* 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved.
* 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z).
=== 2025-03-13 ===
* 14:36 [[gitlab:iccander|@iccander]] was approved.
=== 2025-03-12 ===
* 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z).
* 15:30 [[gitlab:naomi|@naomi]] was approved.
* 15:27 [[gitlab:cobi|@cobi]] was approved.
=== 2025-03-11 ===
* 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z).
=== 2025-03-10 ===
* 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved.
=== 2025-03-09 ===
* 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved.
* 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved.
=== 2025-03-07 ===
* 16:54 [[gitlab:hswan|@hswan]] was approved.
* 14:42 [[gitlab:atitkov|@atitkov]] was approved.
* 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved.
=== 2025-03-06 ===
* 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z).
=== 2025-03-05 ===
* 07:33 [[gitlab:monx9494|@monx9494]] was approved.
=== 2025-03-02 ===
* 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z).
=== 2025-03-01 ===
* 19:15 [[gitlab:izno|@izno]] was approved.
* 12:45 [[gitlab:nyerho|@nyerho]] was approved.
=== 2025-02-28 ===
* 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved.
* 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z).
* 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z).
=== 2025-02-27 ===
* 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z).
=== 2025-02-24 ===
* 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved.
* 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved.
* 19:06 [[gitlab:dhardy|@dhardy]] was approved.
=== 2025-02-22 ===
* 19:27 [[gitlab:owuh|@owuh]] was approved.
=== 2025-02-19 ===
* 16:06 [[gitlab:artemkloko|@artemkloko]] was approved.
* 13:03 [[gitlab:jgafnea|@jgafnea]] was approved.
=== 2025-02-17 ===
* 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved.
=== 2025-02-16 ===
* 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z).
=== 2025-02-15 ===
* 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved.
=== 2025-02-14 ===
* 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z).
=== 2025-02-13 ===
* 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved.
* 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved.
=== 2025-02-12 ===
* 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z).
* 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved.
=== 2025-02-11 ===
* 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved.
=== 2025-02-09 ===
* 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z).
=== 2025-02-07 ===
* 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z).
=== 2025-02-06 ===
* 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z).
* 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved.
=== 2025-02-05 ===
* 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z).
=== 2025-02-03 ===
* 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z).
* 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z).
=== 2025-02-02 ===
* 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z).
=== 2025-01-31 ===
* 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z).
* 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z).
* 00:42 [[gitlab:farid|@farid]] was approved.
=== 2025-01-27 ===
* 08:15 [[gitlab:eliza189|@eliza189]] was approved.
=== 2025-01-25 ===
* 09:51 [[gitlab:pamputt|@pamputt]] was approved.
=== 2025-01-23 ===
* 14:30 [[gitlab:lubianat|@lubianat]] was approved.
* 11:45 [[gitlab:bootsa|@bootsa]] was approved.
=== 2025-01-21 ===
* 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z).
* 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z).
* 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z).
* 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z).
* 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z).
* 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z).
* 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z).
* 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z).
* 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z).
* 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z).
* 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z).
* 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z).
* 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z).
* 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z).
* 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z).
* 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z).
* 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z).
* 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z).
* 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z).
* 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z).
* 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z).
* 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z).
* 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z).
* 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z).
* 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z).
* 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z).
* 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z).
* 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z).
* 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z).
* 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z).
* 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z).
* 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z).
* 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z).
* 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z).
* 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z).
* 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z).
=== 2025-01-20 ===
* 14:33 [[gitlab:your1|@your1]] was approved.
=== 2025-01-18 ===
* 10:09 [[gitlab:galrach600|@galrach600]] was approved.
* 02:51 [[gitlab:blankeclair|@blankeclair]] was approved.
=== 2025-01-17 ===
* 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved.
=== 2025-01-15 ===
* 17:12 [[gitlab:smartse|@smartse]] was approved.
=== 2025-01-14 ===
* 17:03 [[gitlab:naorleizer|@naorleizer]] was approved.
=== 2025-01-13 ===
* 02:45 [[gitlab:wolf20482|@wolf20482]] was approved.
=== 2025-01-12 ===
* 17:45 [[gitlab:tamzin|@tamzin]] was approved.
=== 2025-01-11 ===
* 15:24 [[gitlab:bargioni|@bargioni]] was approved.
* 14:30 [[gitlab:salelya|@salelya]] was approved.
* 10:15 [[gitlab:malakatshy|@malakatshy]] was approved.
* 05:21 [[gitlab:newmcpee|@newmcpee]] was approved.
=== 2025-01-09 ===
* 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved.
=== 2025-01-08 ===
* 16:21 [[gitlab:ukrface|@ukrface]] was approved.
=== 2024-12-28 ===
* 03:27 [[gitlab:twonum|@twonum]] was approved.
=== 2024-12-25 ===
* 06:09 [[gitlab:harsv567|@harsv567]] was approved.
=== 2024-12-21 ===
* 11:24 [[gitlab:amutha2002|@amutha2002]] was approved.
=== 2024-12-20 ===
* 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved.
* 10:00 [[gitlab:ro-shines|@ro-shines]] was approved.
* 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved.
=== 2024-12-18 ===
* 14:45 [[gitlab:soylacarli|@soylacarli]] was approved.
=== 2024-12-16 ===
* 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved.
=== 2024-12-15 ===
* 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved.
=== 2024-12-13 ===
* 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved.
=== 2024-12-10 ===
* 06:39 [[gitlab:ginaan|@ginaan]] was approved.
=== 2024-12-09 ===
* 05:45 [[gitlab:kallinavya|@kallinavya]] was approved.
* 00:54 [[gitlab:viserion-7|@viserion-7]] was approved.
=== 2024-12-08 ===
* 17:27 [[gitlab:wargo|@wargo]] was approved.
=== 2024-12-05 ===
* 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved.
=== 2024-12-02 ===
* 21:21 [[gitlab:a930913|@a930913]] was approved.
=== 2024-12-01 ===
* 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved.
=== 2024-11-21 ===
* 13:45 [[gitlab:sascha|@sascha]] was approved.
=== 2024-11-19 ===
* 16:36 [[gitlab:jly|@jly]] was approved.
=== 2024-11-15 ===
* 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved.
=== 2024-11-14 ===
* 14:15 [[gitlab:stimoroll|@stimoroll]] was approved.
=== 2024-11-09 ===
* 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved.
=== 2024-11-07 ===
* 19:15 [[gitlab:zulf|@zulf]] was approved.
* 05:33 [[gitlab:hassanamin|@hassanamin]] was approved.
=== 2024-11-06 ===
* 19:39 [[gitlab:daniuu|@daniuu]] was approved.
* 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved.
=== 2024-10-09 ===
* 14:45 [[gitlab:jtweed|@jtweed]] was approved.
* 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved.
* 09:06 [[gitlab:wikibayer|@wikibayer]] was approved.
=== 2024-10-06 ===
* 10:27 [[gitlab:keerthan16|@keerthan16]] was approved.
=== 2024-10-04 ===
* 07:45 [[gitlab:hakimi97|@hakimi97]] was approved.
=== 2024-09-30 ===
* 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved.
=== 2024-09-28 ===
* 17:30 [[gitlab:webrunner95|@webrunner95]] was approved.
=== 2024-09-18 ===
* 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved.
=== 2024-09-14 ===
* 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved.
=== 2024-09-06 ===
* 08:48 [[gitlab:mickabarber|@mickabarber]] was approved.
=== 2024-08-27 ===
* 17:36 [[gitlab:edgars|@edgars]] was approved.
=== 2024-08-22 ===
* 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved.
=== 2024-08-14 ===
* 19:21 [[gitlab:jfk|@jfk]] was approved.
=== 2024-08-13 ===
* 17:57 [[gitlab:daxserver|@daxserver]] was approved.
=== 2024-08-11 ===
* 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved.
=== 2024-08-10 ===
* 08:42 [[gitlab:ashig|@ashig]] was approved.
=== 2024-08-09 ===
* 14:09 [[gitlab:masssly|@masssly]] was approved.
=== 2024-08-05 ===
* 22:15 [[gitlab:mrtortue|@mrtortue]] was approved.
=== 2024-08-02 ===
* 16:21 [[gitlab:dsantini|@dsantini]] was approved.
=== 2024-07-31 ===
* 11:54 [[gitlab:cptviraj|@cptviraj]] was approved.
=== 2024-07-30 ===
* 19:09 [[gitlab:iniquity|@iniquity]] was approved.
* 10:00 [[gitlab:collins|@collins]] was approved.
=== 2024-07-27 ===
* 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved.
=== 2024-07-25 ===
* 12:36 [[gitlab:mszabo|@mszabo]] was approved.
* 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved.
=== 2024-07-24 ===
* 08:05 [[gitlab:dragoniez|@dragoniez]] was approved.
=== 2024-07-23 ===
* 06:54 [[gitlab:mirji|@mirji]] was approved.
=== 2024-07-16 ===
* 10:00 [[gitlab:lakejason0|@lakejason0]] was approved.
=== 2024-07-12 ===
* 11:33 [[gitlab:cn|@cn]] was approved.
* 08:12 [[gitlab:unchampignon|@unchampignon]] was approved.
=== 2024-07-07 ===
* 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved.
* 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved.
=== 2024-07-06 ===
* 11:18 [[gitlab:dibya|@dibya]] was approved.
* 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved.
=== 2024-07-05 ===
* 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved.
=== 2024-07-02 ===
* 19:00 [[gitlab:ebrahim|@ebrahim]] was approved.
=== 2024-07-01 ===
* 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved.
* 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved.
=== 2024-06-30 ===
* 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved.
* 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved.
* 06:30 [[gitlab:rihaan180|@rihaan180]] was approved.
=== 2024-06-27 ===
* 17:36 [[gitlab:driedmueller|@driedmueller]] was approved.
=== 2024-06-19 ===
* 12:57 [[gitlab:audreypenven|@audreypenven]] was approved.
=== 2024-06-16 ===
* 01:18 [[gitlab:roysmith|@roysmith]] was approved.
=== 2024-06-08 ===
* 02:45 [[gitlab:jleedev|@jleedev]] was approved.
=== 2024-06-03 ===
* 13:57 [[gitlab:afeder|@afeder]] was approved.
=== 2024-06-01 ===
* 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved.
=== 2024-05-30 ===
* 16:42 [[gitlab:krlsca|@krlsca]] was approved.
=== 2024-05-28 ===
* 11:24 [[gitlab:rickijay|@rickijay]] was approved.
=== 2024-05-26 ===
* 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved.
=== 2024-05-25 ===
* 07:24 [[gitlab:jony|@jony]] was approved.
=== 2024-05-23 ===
* 08:45 [[gitlab:lepticed7|@lepticed7]] was approved.
=== 2024-05-22 ===
* 20:42 [[gitlab:echecs|@echecs]] was approved.
=== 2024-05-21 ===
* 13:33 [[gitlab:mbs|@mbs]] was approved.
=== 2024-05-19 ===
* 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved.
=== 2024-05-18 ===
* 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved.
=== 2024-05-17 ===
* 08:54 [[gitlab:grapedog|@grapedog]] was approved.
=== 2024-05-08 ===
* 19:42 [[gitlab:kelhurd|@kelhurd]] was approved.
* 19:06 [[gitlab:khurd|@khurd]] was approved.
=== 2024-05-06 ===
* 19:48 [[gitlab:j3j5|@j3j5]] was approved.
* 12:06 [[gitlab:tk-999|@tk-999]] was approved.
=== 2024-05-05 ===
* 22:09 [[gitlab:pppery|@pppery]] was approved.
* 20:33 [[gitlab:sakretsu|@sakretsu]] was approved.
* 12:12 [[gitlab:waterquark|@waterquark]] was approved.
=== 2024-05-04 ===
* 09:03 [[gitlab:multichill|@multichill]] was approved.
* 07:42 [[gitlab:abaris|@abaris]] was approved.
=== 2024-05-03 ===
* 14:57 [[gitlab:maurusian|@maurusian]] was approved.
=== 2024-04-24 ===
* 05:48 [[gitlab:wolfinux|@wolfinux]] was approved.
=== 2024-04-23 ===
* 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved.
=== 2024-04-21 ===
* 06:51 [[gitlab:alon|@alon]] was approved.
=== 2024-04-17 ===
* 23:33 [[gitlab:derenrich|@derenrich]] was approved.
=== 2024-04-16 ===
* 17:18 [[gitlab:valcio|@valcio]] was approved.
=== 2024-04-14 ===
* 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved.
=== 2024-04-06 ===
* 12:48 [[gitlab:theprotonade|@theprotonade]] was approved.
=== 2024-04-02 ===
* 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved.
=== 2024-03-30 ===
* 13:36 [[gitlab:lpintscher|@lpintscher]] was approved.
=== 2024-03-26 ===
* 17:09 [[gitlab:eenabulele|@eenabulele]] was approved.
=== 2024-03-25 ===
* 14:27 [[gitlab:tuukka|@tuukka]] was approved.
=== 2024-03-24 ===
* 12:24 [[gitlab:firefly|@firefly]] was approved.
=== 2024-03-21 ===
* 19:33 [[gitlab:universal-omega|@universal-omega]] was approved.
=== 2024-03-17 ===
* 10:36 [[gitlab:bisel91|@bisel91]] was approved.
=== 2024-03-16 ===
* 10:09 [[gitlab:delord|@delord]] was approved.
* 00:42 [[gitlab:athulvis1|@athulvis1]] was approved.
=== 2024-03-15 ===
* 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved.
* 08:30 [[gitlab:peachey88|@peachey88]] was approved.
* 06:51 [[gitlab:derick|@derick]] was approved.
=== 2024-03-12 ===
* 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved.
=== 2024-03-06 ===
* 13:21 [[gitlab:desianabae1|@desianabae1]] was approved.
=== 2024-03-05 ===
* 19:21 [[gitlab:ep1c|@ep1c]] was approved.
* 16:33 [[gitlab:jasmine|@jasmine]] was approved.
=== 2024-03-02 ===
* 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved.
=== 2024-02-29 ===
* 23:18 [[gitlab:arandomname123|@arandomname123]] was approved.
* 18:03 [[gitlab:baba|@baba]] was approved.
* 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved.
* 03:09 [[gitlab:sds|@sds]] was approved.
=== 2024-02-27 ===
* 23:33 [[gitlab:lofhi|@lofhi]] was approved.
=== 2024-02-15 ===
* 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved.
=== 2024-02-14 ===
* 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved.
=== 2024-02-13 ===
* 13:06 [[gitlab:dringsim|@dringsim]] was approved.
=== 2024-02-12 ===
* 17:36 [[gitlab:haak|@haak]] was approved.
=== 2024-02-05 ===
* 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved.
* 17:14 [[gitlab:ahecht|@ahecht]] was approved.
=== 2024-02-01 ===
* 09:27 [[gitlab:arinaigum|@arinaigum]] was approved.
* 00:15 [[gitlab:jas42|@jas42]] was approved.
* 00:15 [[gitlab:edhu|@edhu]] was approved.
* 00:15 [[gitlab:marnanel|@marnanel]] was approved.
* 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved.
* 00:15 [[gitlab:amasotti|@amasotti]] was approved.
* 00:15 [[gitlab:deni|@deni]] was approved.
* 00:15 [[gitlab:cyber|@cyber]] was approved.
* 00:15 [[gitlab:saroj|@saroj]] was approved.
=== 2024-01-29 ===
* 21:42 [[gitlab:rgupta|@rgupta]] was approved.
=== 2024-01-07 ===
* 09:48 [[gitlab:lutrome|@lutrome]] was approved.
=== 2024-01-05 ===
* 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved.
* 02:51 [[gitlab:braunobruno|@braunobruno]] was approved.
* 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved.
* 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved.
=== 2024-01-03 ===
* 14:45 [[gitlab:gabina|@gabina]] was approved.
=== 2024-01-02 ===
* 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved.
=== 2023-12-23 ===
* 00:33 [[gitlab:aram|@aram]] was approved.
=== 2023-12-22 ===
* 16:24 [[gitlab:elpitareio|@elpitareio]] was approved.
=== 2023-12-21 ===
* 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved.
* 00:43 [[gitlab:ederporto|@ederporto]] was approved.
* 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved.
* 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved.
* 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved.
* 00:43 [[gitlab:air7538|@air7538]] was approved.
* 00:43 [[gitlab:anzx|@anzx]] was approved.
* 00:43 [[gitlab:tekask1903|@tekask1903]] was approved.
* 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved.
* 00:42 [[gitlab:mpaa|@mpaa]] was approved.
* 00:42 [[gitlab:kutay|@kutay]] was approved.
* 00:42 [[gitlab:wattmto|@wattmto]] was approved.
9s0l0ej7y0zx3iakolq9965tf3tt25z
Help:Toolforge/My first Python tool
12
460669
2461129
2460066
2026-09-26T20:50:11Z
SSapaty (WMF)
28955
collapse block of app.py code sample
2461129
wikitext
text/x-wiki
= My first Python tool =
== Overview ==
Python webservices are used by many existing tools on Toolforge. [[w:Python (programming language)|Python]] is a general-purpose programming language commonly used for web development, automation, data processing, and many other tasks.
This tutorial is designed to get a sample Python application deployed to Toolforge using [[Help:Toolforge/Deploy your tool|Push-to-Deploy]] as quickly as possible. The application uses [[w:Flask (web framework)|Flask]] and runs on Toolforge using the Build Service.
Then we'll use the database credentials provided by Toolforge to extract the list of available databases from the Wiki Replicas.
The guide will teach you how to:
* Create a new tool
* Run a Python Flask webservice on [[w:Kubernetes|Kubernetes]]
* Deploy the application using [[Help:Toolforge/Deploy your tool|Push-to-Deploy]]
* Connect to the Wiki Replicas from Python
* Run a query and extract information
== Getting started ==
=== Prerequisites ===
==== Skills ====
* Basic knowledge of [[w:Python (programming language)|Python]]
* Basic knowledge of [[w:Secure Shell|SSH]]
* Basic knowledge of the [[w:Unix shell|Unix command line]]
* Basic knowledge of [[w:Git|Git]]
==== Accounts ====
* [[Help:Toolforge/Quickstart|A Toolforge account]]
== Step-by-step guide ==
=== Step 1: Create a new tool account ===
# Follow the [[Help:Toolforge/Quickstart|Toolforge quickstart]] guide to create a Toolforge tool and SSH into Toolforge.
#* For the examples in this tutorial, <code>sample-python-ptd-app</code> is used to indicate places where your unique tool name is used in another command.
# Make sure to create a Git repository for the tool. You can get one like this:
## Log into the [https://toolsadmin.wikimedia.org/ Toolforge admin page].
## Select your tool.
## On the left side panel, under <code>Git repositories</code>, click <code>create repository</code>.
## Copy the URL in the <code>Clone</code> section.
### There is a private URL that we will use to clone the repository locally, starting with <code>git</code>: <code>git@gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app.git</code>
### There is also a public URL that Toolforge will use to build the application, starting with <code>https</code>: <code>https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app.git</code>
=== Step 2: Create a basic Python Flask webservice ===
What is Flask?
[[w:Flask (web framework)|Flask]] is a lightweight Python web framework that can be used to create web applications and APIs.
==== How to create a basic Python Flask webservice ====
'''Clone your tool Git repository'''
You will have to clone the tool repository to be able to add code to it. On your local computer, with Git installed, run:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~$ git clone git@gitlab.wikimedia.org:toolforge-repos/sample-python-ptd-app.git
laptop:~$ cd sample-python-ptd-app
}}
That will create a folder called <code>sample-python-ptd-app</code>. We are going to put the code in that folder.
'''Add the Python dependencies'''
{{Note|Use a file named <code>requirements.txt</code> to keep track of the Python dependencies required by the application.}}
Create <code>requirements.txt</code>:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ cat > requirements.txt << EOF
Flask
gunicorn
PyMySQL
EOF
}}
'''Create a "Hello World!" Python application'''
Create a file named <code>app.py</code>:
{{Codesample|lang=python|scheme=light|name=app.py|code=
from flask import Flask
app = Flask(__name__)
@app.route("/")
def index():
return "<p>Hello World!</p>"
}}
{{Note|Code on Toolforge must always be licensed under an [[w:Open-source license|Open Source Initiative (OSI) approved license]]. See [[Help:Toolforge/Right to fork policy|Right to fork policy]] for more information on this Toolforge policy.}}
'''Create the Procfile'''
The <code>Procfile</code> defines the command that the Build Service should use to start the web application.
Create it with:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ cat > Procfile << EOF
web: gunicorn --bind 0.0.0.0:\$PORT app:app
EOF
}}
The <code>web</code> entry defines the command used to start the Flask application. Gunicorn runs the application and listens on the port provided by Toolforge through the <code>PORT</code> environment variable.
'''Create the Toolforge configuration file'''
Create <code>toolforge.yaml</code>:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ cat > toolforge.yaml << EOF
# yaml-language-server: \$schema=https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/raw/main/openapi/tool-config-schema.json
source_url: https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app/-/raw/main/toolforge.yaml?ref_type=heads
components:
webservice:
build:
repository: https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app
ref: main
run:
command: web
publish: /
port: 8000
EOF
}}
If you don't want your webservice to be reachable from the internet, remove the <code>publish</code> setting. If you are deploying a process without a port, you can remove <code>port</code> too. See [[Help:Toolforge/Deploy your tool|Deploy your tool]] for more information.
'''Commit your changes and push'''
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ git add .
laptop:~/sample-python-ptd-app$ git commit -m "First commit"
laptop:~/sample-python-ptd-app$ git push origin main
}}
'''Initialize the tool configuration'''
Now SSH to <code>login.toolforge.org</code> and become your tool:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ ssh login.toolforge.org
user@tools-sgebastion-XX:~$ become sample-python-ptd-app
}}
Create the tool configuration using the public copy of <code>toolforge.yaml</code>:
{{Codesample|lang=shell-session|scheme=light|code=
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ curl 'https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app/-/raw/main/toolforge.yaml?ref_type=heads' {{!}} toolforge components config create
}}
You should see a message indicating that the configuration was updated successfully.
{{Note|Push-to-Deploy is currently a beta feature of Toolforge, so the command may also display a warning indicating that you are using a beta feature.}}
'''Create the first deployment'''
Create a deployment:
{{Codesample|lang=shell-session|scheme=light|code=
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment create
}}
Push-to-Deploy will build the application from your Git repository and deploy the webservice.
{{Note|The repository in <code>toolforge.yaml</code> must be publicly accessible so that Toolforge can clone and build it.}}
'''Wait for the deployment to finish'''
Check the deployment status:
{{Codesample|lang=shell-session|scheme=light|code=
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment show
}}
Wait for the deployment status to be <code>ok(Succeeded)</code>.
Once the deployment has finished, navigate to:
{{Codesample|lang=shell-session|scheme=light|code=
https://sample-python-ptd-app.toolforge.org
}}
You should see:
<syntaxhighlight lang="text">
Hello World!
</syntaxhighlight>
It might take a couple of minutes for the application to become reachable.
'''Automate deployment on push'''
You can follow [[Help:Toolforge/Deploy your tool#Automatically deploy on git push|this guide]] to configure and automate deployment of your tool whenever you push or merge changes to the main branch.
==== Connecting to the replicas ====
Toolforge provides database credentials to your application through environment variables. We will use these credentials to connect to the special <code>meta_p</code> database and retrieve the names of the available Wiki Replica databases.
First, update <code>requirements.txt</code> if <code>PyMySQL</code> is not already present:
<syntaxhighlight lang="text">
Flask
gunicorn
PyMySQL
</syntaxhighlight>
Now update <code>app.py</code>:
{{Collapse top|app.y}}
{{Codesample|lang=python|scheme=light|name=app.py|code=
from flask import Flask
import os
import pymysql
app = Flask(__name__)
def get_db_connection():
return pymysql.connect(
host="meta.web.db.svc.wikimedia.cloud",
user=os.environ["TOOL_REPLICA_USER"],
password=os.environ["TOOL_REPLICA_PASSWORD"],
database="meta_p",
cursorclass=pymysql.cursors.DictCursor,
)
def get_db_names():
connection = get_db_connection()
try:
with connection.cursor() as cursor:
cursor.execute("SELECT dbname FROM wiki;")
return cursor.fetchall()
finally:
connection.close()
@app.route("/")
def index():
databases = get_db_names()
rows = "".join(
f"<tr><td>{row['dbname']}</td></tr>"
for row in databases
)
return f"""
<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Database names</title>
</head>
<body>
<p>
Sample Python application accessing the
<a href="https://wikitech.wikimedia.org/wiki/Help:Toolforge/Database">
Wiki Replicas
</a>
special <code>meta_p</code> database.
</p>
<table>
<thead>
<tr>
<th>Database name</th>
</tr>
</thead>
<tbody>
{rows}
</tbody>
</table>
</body>
</html>
"""
}}
{{Collapse bottom}}
The application connects to <code>meta.web.db.svc.wikimedia.cloud</code> using the <code>TOOL_REPLICA_USER</code> and <code>TOOL_REPLICA_PASSWORD</code> environment variables provided by Toolforge.
The query:
{{Codesample|lang=sql|scheme=light|code=
SELECT dbname FROM wiki;
}}
retrieves the database names from the special <code>meta_p</code> database.
'''Commit, push, and redeploy'''
Commit and push your changes:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ git add .
laptop:~/sample-python-ptd-app$ git commit -m "Query meta_p database"
laptop:~/sample-python-ptd-app$ git push origin main
}}
Then create a new deployment:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ ssh login.toolforge.org
user@tools-sgebastion-XX:~$ become sample-python-ptd-app
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment create
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment show
}}
Wait for the deployment to finish and refresh the application in your browser.
You should now see a list of Wiki Replica database names.
== Configure Push-to-Deploy ==
Push-to-Deploy allows a CI runner to trigger a Components Service deployment after changes are pushed to your Git repository.
Below are the steps for a basic configuration. Visit [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] for more details.
{{Collapse top|Configure Push-to-Deploy}}
== Configure Push-to-Deploy ==
To enable Push-to-Deploy, configure your Git repository's CI system to trigger a Components deployment whenever changes are pushed to the repository.
The following is a minimal example for Wikimedia GitLab. For additional options and details, see [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]].
=== Create a deploy token ===
While logged in as your tool account, create a deploy token:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deploy-token create
}}
The command displays a token. Treat this token as a secret. Do not commit it to your repository.
=== Add the deploy token to GitLab ===
In your GitLab repository, go to:
<code>Settings → CI/CD → Variables</code>
Create a variable named:
{{Codesample|lang=text|scheme=light|code=
TOOLS_DEPLOY_TOKEN
}}
Set its value to the deploy token created in the previous step.
For security, make the variable hidden, protected, and masked where appropriate.
=== Add a GitLab CI configuration ===
Create a <code>.gitlab-ci.yml</code> file in the repository containing:
{{Codesample|lang=yaml|scheme=light|name=.gitlab-ci.yml|code=
include:
- project: "repos/cloud/cicd/gitlab-ci"
file: "toolforge-cd/deploy-to-toolforge.yaml"
}}
The included Toolforge CI configuration handles triggering the Components deployment.
This minimal configuration assumes that the GitLab project name and Toolforge tool name are the same. If they are different, or if you need to customize the deployment process, see [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]].
Commit and push <code>.gitlab-ci.yml</code> to the repository.
When a commit is pushed or merged to the default branch, GitLab CI will use the deploy token to request a new Components deployment for <code>komla-test3</code>.
When changes are pushed to the repository, GitLab CI will use the deploy token to trigger a Components deployment.
At this point, Push-to-Deploy is configured.
{{Collapse bottom}}
=== Test Push-to-Deploy ===
After configuring your CI runner, make a small change to <code>app.py</code>.
For example, change the title from `Database names` to `All Database Names:
{{Codesample|lang=yaml|scheme=light|name=app.py|code=
<title>All Database Names</title>
}}
Commit and push the change:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add app.py
$ git commit -m "Expand database name"
$ git push
}}
Do not manually create a deployment.
The push should trigger your CI runner, which creates a Components Service deployment automatically.
Check your repository's CI pipeline and verify that it completes successfully.
You can then inspect the deployment from Toolforge:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
For the scheduled example, verify that the new schedule has been applied:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
If the deployment was triggered automatically and the new change was applied, Push-to-Deploy is working.
==== Notes ====
You can see the code used in this example in the <code>sample-python-ptd-app</code> repository.
== Troubleshooting ==
See [[Help:Toolforge/Build Service#Troubleshooting|Toolforge Build Service troubleshooting]].
You can also inspect the status of your Push-to-Deploy deployments with the following:
{{Codesample|lang=shell-session|scheme=light|code=
toolforge components deployment list
toolforge components deployment show
toolforge jobs logs webservice #view logs
}}
== See also ==
* [[Help:Toolforge/Deploy your tool|Deploy your tool with Push-to-Deploy]]
* [[Help:Toolforge/Build Service|Toolforge Build Service]]
* [[Help:Toolforge/Database|Toolforge databases]]
* [[Help:Toolforge/Quickstart|Toolforge quickstart]]
* [[Help:Toolforge/My first PHP tool|My first PHP tool]]
== Communication and support ==
Support and administration of the WMCS resources is provided by Wikimedia Foundation staff and Wikimedia movement volunteers. Please reach out with questions and join the conversation:
'''Discuss and receive general support'''
* Chat in real time in the [[Help:IRC|IRC channel]] <code>#wikimedia-cloud</code> or the bridged Telegram group.
* Discuss via email after you have subscribed to the [[mail:cloud|cloud@]] mailing list.
'''Stay aware of critical changes and plans'''
* Subscribe to the [[mail:cloud-announce|cloud-announce@]] mailing list. Messages are also mirrored to the [[mail:cloud|cloud@]] list.
'''Track work tasks and report bugs'''
Use a subproject of the [[phab:tag/cloud-services/|Cloud-Services]] Phabricator project to track confirmed bug reports and feature requests about the Cloud Services infrastructure itself.
{{:Help:Cloud Services communication}}
[[Category:Toolforge]]
[[Category:Tutorials|Python]]
[[Category:How-to-guide|Python]]
[[Category:Python]]
r03dbhbwprtbetgflfnrasuaxd7buij
2461130
2461129
2026-09-26T20:51:38Z
SSapaty (WMF)
28955
2461130
wikitext
text/x-wiki
= My first Python tool =
== Overview ==
Python webservices are used by many existing tools on Toolforge. [[w:Python (programming language)|Python]] is a general-purpose programming language commonly used for web development, automation, data processing, and many other tasks.
This tutorial is designed to get a sample Python application deployed to Toolforge using [[Help:Toolforge/Deploy your tool|Push-to-Deploy]] as quickly as possible. The application uses [[w:Flask (web framework)|Flask]] and runs on Toolforge using the Build Service.
Then we'll use the database credentials provided by Toolforge to extract the list of available databases from the Wiki Replicas.
The guide will teach you how to:
* Create a new tool
* Run a Python Flask webservice on [[w:Kubernetes|Kubernetes]]
* Deploy the application using [[Help:Toolforge/Deploy your tool|Push-to-Deploy]]
* Connect to the Wiki Replicas from Python
* Run a query and extract information
== Getting started ==
=== Prerequisites ===
==== Skills ====
* Basic knowledge of [[w:Python (programming language)|Python]]
* Basic knowledge of [[w:Secure Shell|SSH]]
* Basic knowledge of the [[w:Unix shell|Unix command line]]
* Basic knowledge of [[w:Git|Git]]
==== Accounts ====
* [[Help:Toolforge/Quickstart|A Toolforge account]]
== Step-by-step guide ==
=== Step 1: Create a new tool account ===
# Follow the [[Help:Toolforge/Quickstart|Toolforge quickstart]] guide to create a Toolforge tool and SSH into Toolforge.
#* For the examples in this tutorial, <code>sample-python-ptd-app</code> is used to indicate places where your unique tool name is used in another command.
# Make sure to create a Git repository for the tool. You can get one like this:
## Log into the [https://toolsadmin.wikimedia.org/ Toolforge admin page].
## Select your tool.
## On the left side panel, under <code>Git repositories</code>, click <code>create repository</code>.
## Copy the URL in the <code>Clone</code> section.
### There is a private URL that we will use to clone the repository locally, starting with <code>git</code>: <code>git@gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app.git</code>
### There is also a public URL that Toolforge will use to build the application, starting with <code>https</code>: <code>https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app.git</code>
=== Step 2: Create a basic Python Flask webservice ===
What is Flask?
[[w:Flask (web framework)|Flask]] is a lightweight Python web framework that can be used to create web applications and APIs.
==== How to create a basic Python Flask webservice ====
'''Clone your tool Git repository'''
You will have to clone the tool repository to be able to add code to it. On your local computer, with Git installed, run:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~$ git clone git@gitlab.wikimedia.org:toolforge-repos/sample-python-ptd-app.git
laptop:~$ cd sample-python-ptd-app
}}
That will create a folder called <code>sample-python-ptd-app</code>. We are going to put the code in that folder.
'''Add the Python dependencies'''
{{Note|Use a file named <code>requirements.txt</code> to keep track of the Python dependencies required by the application.}}
Create <code>requirements.txt</code>:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ cat > requirements.txt << EOF
Flask
gunicorn
PyMySQL
EOF
}}
'''Create a "Hello World!" Python application'''
Create a file named <code>app.py</code>:
{{Codesample|lang=python|scheme=light|name=app.py|code=
from flask import Flask
app = Flask(__name__)
@app.route("/")
def index():
return "<p>Hello World!</p>"
}}
{{Note|Code on Toolforge must always be licensed under an [[w:Open-source license|Open Source Initiative (OSI) approved license]]. See [[Help:Toolforge/Right to fork policy|Right to fork policy]] for more information on this Toolforge policy.}}
'''Create the Procfile'''
The <code>Procfile</code> defines the command that the Build Service should use to start the web application.
Create it with:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ cat > Procfile << EOF
web: gunicorn --bind 0.0.0.0:\$PORT app:app
EOF
}}
The <code>web</code> entry defines the command used to start the Flask application. Gunicorn runs the application and listens on the port provided by Toolforge through the <code>PORT</code> environment variable.
'''Create the Toolforge configuration file'''
Create <code>toolforge.yaml</code>:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ cat > toolforge.yaml << EOF
# yaml-language-server: \$schema=https://gitlab.wikimedia.org/repos/cloud/toolforge/components-api/-/raw/main/openapi/tool-config-schema.json
source_url: https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app/-/raw/main/toolforge.yaml?ref_type=heads
components:
webservice:
build:
repository: https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app
ref: main
run:
command: web
publish: /
port: 8000
EOF
}}
If you don't want your webservice to be reachable from the internet, remove the <code>publish</code> setting. If you are deploying a process without a port, you can remove <code>port</code> too. See [[Help:Toolforge/Deploy your tool|Deploy your tool]] for more information.
'''Commit your changes and push'''
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ git add .
laptop:~/sample-python-ptd-app$ git commit -m "First commit"
laptop:~/sample-python-ptd-app$ git push origin main
}}
'''Initialize the tool configuration'''
Now SSH to <code>login.toolforge.org</code> and become your tool:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ ssh login.toolforge.org
user@tools-sgebastion-XX:~$ become sample-python-ptd-app
}}
Create the tool configuration using the public copy of <code>toolforge.yaml</code>:
{{Codesample|lang=shell-session|scheme=light|code=
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ curl 'https://gitlab.wikimedia.org/toolforge-repos/sample-python-ptd-app/-/raw/main/toolforge.yaml?ref_type=heads' {{!}} toolforge components config create
}}
You should see a message indicating that the configuration was updated successfully.
{{Note|Push-to-Deploy is currently a beta feature of Toolforge, so the command may also display a warning indicating that you are using a beta feature.}}
'''Create the first deployment'''
Create a deployment:
{{Codesample|lang=shell-session|scheme=light|code=
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment create
}}
Push-to-Deploy will build the application from your Git repository and deploy the webservice.
{{Note|The repository in <code>toolforge.yaml</code> must be publicly accessible so that Toolforge can clone and build it.}}
'''Wait for the deployment to finish'''
Check the deployment status:
{{Codesample|lang=shell-session|scheme=light|code=
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment show
}}
Wait for the deployment status to be <code>ok(Succeeded)</code>.
Once the deployment has finished, navigate to:
{{Codesample|lang=shell-session|scheme=light|code=
https://sample-python-ptd-app.toolforge.org
}}
You should see:
<syntaxhighlight lang="text">
Hello World!
</syntaxhighlight>
It might take a couple of minutes for the application to become reachable.
'''Automate deployment on push'''
You can follow [[Help:Toolforge/Deploy your tool#Automatically deploy on git push|this guide]] to configure and automate deployment of your tool whenever you push or merge changes to the main branch.
==== Connecting to the replicas ====
Toolforge provides database credentials to your application through environment variables. We will use these credentials to connect to the special <code>meta_p</code> database and retrieve the names of the available Wiki Replica databases.
First, update <code>requirements.txt</code> if <code>PyMySQL</code> is not already present:
<syntaxhighlight lang="text">
Flask
gunicorn
PyMySQL
</syntaxhighlight>
Now update <code>app.py</code>:
{{Collapse top|Sample flask code}}
{{Codesample|lang=python|scheme=light|name=app.py|code=
from flask import Flask
import os
import pymysql
app = Flask(__name__)
def get_db_connection():
return pymysql.connect(
host="meta.web.db.svc.wikimedia.cloud",
user=os.environ["TOOL_REPLICA_USER"],
password=os.environ["TOOL_REPLICA_PASSWORD"],
database="meta_p",
cursorclass=pymysql.cursors.DictCursor,
)
def get_db_names():
connection = get_db_connection()
try:
with connection.cursor() as cursor:
cursor.execute("SELECT dbname FROM wiki;")
return cursor.fetchall()
finally:
connection.close()
@app.route("/")
def index():
databases = get_db_names()
rows = "".join(
f"<tr><td>{row['dbname']}</td></tr>"
for row in databases
)
return f"""
<!doctype html>
<html lang="en">
<head>
<meta charset="UTF-8">
<title>Database names</title>
</head>
<body>
<p>
Sample Python application accessing the
<a href="https://wikitech.wikimedia.org/wiki/Help:Toolforge/Database">
Wiki Replicas
</a>
special <code>meta_p</code> database.
</p>
<table>
<thead>
<tr>
<th>Database name</th>
</tr>
</thead>
<tbody>
{rows}
</tbody>
</table>
</body>
</html>
"""
}}
{{Collapse bottom}}
The application connects to <code>meta.web.db.svc.wikimedia.cloud</code> using the <code>TOOL_REPLICA_USER</code> and <code>TOOL_REPLICA_PASSWORD</code> environment variables provided by Toolforge.
The query:
{{Codesample|lang=sql|scheme=light|code=
SELECT dbname FROM wiki;
}}
retrieves the database names from the special <code>meta_p</code> database.
'''Commit, push, and redeploy'''
Commit and push your changes:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ git add .
laptop:~/sample-python-ptd-app$ git commit -m "Query meta_p database"
laptop:~/sample-python-ptd-app$ git push origin main
}}
Then create a new deployment:
{{Codesample|lang=shell-session|scheme=light|code=
laptop:~/sample-python-ptd-app$ ssh login.toolforge.org
user@tools-sgebastion-XX:~$ become sample-python-ptd-app
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment create
tools.sample-python-ptd-app@tools-sgebastion-XX:~$ toolforge components deployment show
}}
Wait for the deployment to finish and refresh the application in your browser.
You should now see a list of Wiki Replica database names.
== Configure Push-to-Deploy ==
Push-to-Deploy allows a CI runner to trigger a Components Service deployment after changes are pushed to your Git repository.
Below are the steps for a basic configuration. Visit [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] for more details.
== Configure Push-to-Deploy ==
To enable Push-to-Deploy, configure your Git repository's CI system to trigger a Components deployment whenever changes are pushed to the repository.
The following is a minimal example for Wikimedia GitLab. For additional options and details, see [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]].
=== Create a deploy token ===
While logged in as your tool account, create a deploy token:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deploy-token create
}}
The command displays a token. Treat this token as a secret. Do not commit it to your repository.
=== Add the deploy token to GitLab ===
In your GitLab repository, go to:
<code>Settings → CI/CD → Variables</code>
Create a variable named:
{{Codesample|lang=text|scheme=light|code=
TOOLS_DEPLOY_TOKEN
}}
Set its value to the deploy token created in the previous step.
For security, make the variable hidden, protected, and masked where appropriate.
=== Add a GitLab CI configuration ===
Create a <code>.gitlab-ci.yml</code> file in the repository containing:
{{Codesample|lang=yaml|scheme=light|name=.gitlab-ci.yml|code=
include:
- project: "repos/cloud/cicd/gitlab-ci"
file: "toolforge-cd/deploy-to-toolforge.yaml"
}}
The included Toolforge CI configuration handles triggering the Components deployment.
This minimal configuration assumes that the GitLab project name and Toolforge tool name are the same. If they are different, or if you need to customize the deployment process, see [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]].
Commit and push <code>.gitlab-ci.yml</code> to the repository.
When a commit is pushed or merged to the default branch, GitLab CI will use the deploy token to request a new Components deployment for <code>komla-test3</code>.
When changes are pushed to the repository, GitLab CI will use the deploy token to trigger a Components deployment.
At this point, Push-to-Deploy is configured.
=== Test Push-to-Deploy ===
After configuring your CI runner, make a small change to <code>app.py</code>.
For example, change the title from `Database names` to `All Database Names:
{{Codesample|lang=yaml|scheme=light|name=app.py|code=
<title>All Database Names</title>
}}
Commit and push the change:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add app.py
$ git commit -m "Expand database name"
$ git push
}}
Do not manually create a deployment.
The push should trigger your CI runner, which creates a Components Service deployment automatically.
Check your repository's CI pipeline and verify that it completes successfully.
You can then inspect the deployment from Toolforge:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
For the scheduled example, verify that the new schedule has been applied:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
If the deployment was triggered automatically and the new change was applied, Push-to-Deploy is working.
==== Notes ====
You can see the code used in this example in the <code>sample-python-ptd-app</code> repository.
== Troubleshooting ==
See [[Help:Toolforge/Build Service#Troubleshooting|Toolforge Build Service troubleshooting]].
You can also inspect the status of your Push-to-Deploy deployments with the following:
{{Codesample|lang=shell-session|scheme=light|code=
toolforge components deployment list
toolforge components deployment show
toolforge jobs logs webservice #view logs
}}
== See also ==
* [[Help:Toolforge/Deploy your tool|Deploy your tool with Push-to-Deploy]]
* [[Help:Toolforge/Build Service|Toolforge Build Service]]
* [[Help:Toolforge/Database|Toolforge databases]]
* [[Help:Toolforge/Quickstart|Toolforge quickstart]]
* [[Help:Toolforge/My first PHP tool|My first PHP tool]]
== Communication and support ==
Support and administration of the WMCS resources is provided by Wikimedia Foundation staff and Wikimedia movement volunteers. Please reach out with questions and join the conversation:
'''Discuss and receive general support'''
* Chat in real time in the [[Help:IRC|IRC channel]] <code>#wikimedia-cloud</code> or the bridged Telegram group.
* Discuss via email after you have subscribed to the [[mail:cloud|cloud@]] mailing list.
'''Stay aware of critical changes and plans'''
* Subscribe to the [[mail:cloud-announce|cloud-announce@]] mailing list. Messages are also mirrored to the [[mail:cloud|cloud@]] list.
'''Track work tasks and report bugs'''
Use a subproject of the [[phab:tag/cloud-services/|Cloud-Services]] Phabricator project to track confirmed bug reports and feature requests about the Cloud Services infrastructure itself.
{{:Help:Cloud Services communication}}
[[Category:Toolforge]]
[[Category:Tutorials|Python]]
[[Category:How-to-guide|Python]]
[[Category:Python]]
4t5v4c0evfzw4d8zwvc6yjvzrjmvu4d
Help:Toolforge/Migrate build service job to push to deploy
12
460726
2461131
2461062
2026-09-26T20:53:29Z
SSapaty (WMF)
28955
Remove collapsed panel for push-to-deploy configuration instructions
2461131
wikitext
text/x-wiki
{{Toolforge nav}}{{Note|Push-to-Deploy is currently a beta feature of Toolforge.}}
This tutorial explains how to migrate an existing Toolforge job that uses an image built with the [[Help:Toolforge/Build Service|Build Service]] to the Components Service, and then configure [[Help:Toolforge/Deploy your tool|Push-to-Deploy]] so that future Git pushes automatically trigger deployments.
The job we'll be migrating is a simple Toolforge job that fetches the most recent changes on English Wikipedia.
If you already use the Build Service together with the Jobs Framework, most of the information needed to create the Components configuration can be generated automatically from your existing job.
This tutorial uses a scheduled job named <code>recent-changes</code> as an example.
The migration has two main stages:
# Migrate the existing Build Service job to the Components Service and verify that the component works.
# Configure your Git repository to trigger Components deployments automatically when you push changes.
The second stage enables Push-to-Deploy.
== Before you begin ==
This tutorial assumes that:
* you already have a working Toolforge job;
* the job uses an image built with the Toolforge Build Service;
* the source code is stored in a publicly accessible Git repository supported by the Build Service;
* you have access to the Toolforge tool account;
* you can commit and push changes to the source repository.
Become your tool before running Toolforge commands in this tutorial:
{{Codesample|lang=shell-session|scheme=light|code=
$ become <TOOL NAME>
}}
Replace <code><TOOL NAME></code> with the name of your Toolforge tool.
== Inspect your existing job ==
First, inspect the current Jobs Framework configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
For example, the <code>recent-changes</code> job looks similar to this:
{{Codesample|lang=text|scheme=light|code=
Job name: recent-changes
Command: ./recent-changes-job.sh
Job type: scheduled: @daily
Image: tool-komla-test3/recent-changes:latest
File log: no
Emails: none
Resources: default
Mounts: all
Retry: no
Timeout: yes: 3600s
Status: Pending for 50m59s
}}
We see that we have one scheduled job named `recent-changes` that runs `daily` and is using the `python3.13` image.
You can also inspect previous builds:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge build list
}}
For the example job, the relevant build points to:
{{Codesample|lang=text|scheme=light|code=
Source repository: https://gitlab.wikimedia.org/toolforge-repos/komla-test3
Image name: tool-komla-test3/recent-changes:latest
Git ref: main
}}
== Generate a Components configuration ==
The Components Service can generate an example configuration from your existing Jobs Framework jobs.
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config generate
}}
{{Warning|The generated configuration is an example. Review and validate it before deploying it.}}
For the <code>recent-changes</code> job, the command generates a configuration similar to:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
config_version: v1beta1
components:
recent-changes:
build:
repository: https://gitlab.wikimedia.org/toolforge-repos/komla-test3
ref: main
run:
command: python recent-changes.py
schedule: '@daily'
timeout: 3600
}}
The generated configuration translates the existing Jobs Framework configuration into a Components Service configuration.
For example:
{| class="wikitable"
! Existing job setting
! Components configuration setting
|-
| Job name
| Component name
|-
| Build Service source repository
| <code>build.repository</code>
|-
| Git branch or ref
| <code>build.ref</code>
|-
| Job command
| <code>run.command</code>
|-
| CPU
| <code>run.cpu</code>
|-
| Memory
| <code>run.memory</code>
|-
| Email notification policy
| <code>run.emails</code>
|-
| File logging
| <code>run.filelog</code>
|-
| Filesystem mount
| <code>run.mount</code>
|-
| Retry count
| <code>run.retry</code>
|-
| Job timeout
| <code>run.timeout</code>
|-
| Cron schedule
| <code>run.schedule</code>
|-
| Job type
| <code>component_type</code>
|}
== Review the generated configuration ==
Do not deploy the generated configuration without reviewing it.
Check that at least the following values match your existing deployment:
* source repository;
* Git branch or ref;
* command;
* component type;
* CPU and memory;
* schedule, for scheduled jobs;
* retry and timeout settings;
* email notification settings;
* filesystem mount settings.
You should also verify that the generated component represents the job you intend to migrate.
For example:
{{Codesample|lang=yaml|scheme=light|code=
components:
recent-changes:
}}
corresponds to the existing Jobs Framework job named <code>recent-changes</code>.
If your tool contains several jobs, <code>toolforge components config generate</code> may generate configuration for multiple components. Review each generated component before continuing.
== Add the configuration to your source repository ==
The Components configuration can be hosted anywhere that is publicly accessible.
For this tutorial, save the generated configuration as <code>toolforge.yaml</code> in the same Git repository as the source code. Keeping the source code and deployment configuration together makes it easier to track changes to both in Git.
You can use any publicly accessible Git repository. For Toolforge projects, we recommend using a repository in the Wikimedia GitLab <code>toolforge-repos</code> namespace.
For example:
{{Codesample|lang=text|scheme=light|code=
my-project/
├── Procfile
├── requirements.txt
├── source-files...
└── toolforge.yaml
}}
Commit and push the file:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add toolforge.yaml
$ git commit -m "Add Toolforge Components configuration"
$ git push
}}
The repository and ref in <code>toolforge.yaml</code> should point to the source you want Toolforge to build.
== Create the Components configuration ==
Once <code>toolforge.yaml</code> is publicly accessible, create the Components Service configuration from the copy stored in the repository.
For example:
{{Codesample|lang=shell-session|scheme=light|code=
$ curl 'https://gitlab.wikimedia.org/toolforge-repos/komla-test3/-/raw/main/toolforge.yaml?ref_type=heads' {{!}} toolforge components config create
}}
You should see information saying configuration created or updated successfully.
Replace the URL with the public raw URL for your <code>toolforge.yaml</code> file.
This avoids having to clone the repository or download <code>toolforge.yaml</code> on the Toolforge bastion.
Inspect the stored configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
Review the output and make sure it matches the configuration you committed.
== Test the migration with a manual deployment ==
Before configuring Push-to-Deploy, create a deployment manually to verify that the Components Service can successfully manage the workload.
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create
$ toolforge components deployment show #to see deployment status
}}
{{Note|This is a manual Components Service deployment. Running <code>toolforge components deployment create</code> yourself is not Push-to-Deploy. Push-to-Deploy will be configured later in this tutorial so that your CI runner creates deployments automatically after you push changes to Git.}}
The Components Service will use the configuration to build and deploy the component.
You can optionally attach a description:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --description "Test recent-changes migration"
}}
Normally, the Components Service can reuse an existing build when the source has not changed.
To force a new build, use:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-build
}}
To force the component to run again even when its configuration has not changed, use:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-run
}}
These options are normally unnecessary for the initial migration.
{{Note|When the first Components deployment runs, Components Service builds the configured image and creates or updates the corresponding Jobs Framework job. If a job with the same name already exists, as in this migration, the existing job can be updated in place.}}
== Verify the migrated component ==
Inspect your Components deployments:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
You can also inspect the current Components configuration:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
For a scheduled component, verify that the resulting job exists:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
Compare the resulting job with the configuration you reviewed earlier.
For the example, you would expect to see a job similar to:
{{Codesample|lang=text|scheme=light|code=
Job name: recent-changes
Command: python recent-changes.py
Job type: scheduled: @daily
}}
Also verify that the job behaves as expected when its next scheduled execution occurs.
At this point, you have migrated the workload to the Components Service and verified that it works. The next step is to configure your Git repository so that future deployments are triggered automatically.
== Configure Push-to-Deploy ==
Push-to-Deploy allows a CI runner to trigger a Components Service deployment after changes are pushed to your Git repository.
Below are the steps for a basic configuration. Visit [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] for more details.
ga
The tool is now managed through Components Service, but deployments still need to be started manually.
To enable Push-to-Deploy, configure your Git repository's CI system to trigger a Components deployment whenever changes are pushed to the repository.
The following is a minimal example for Wikimedia GitLab. For additional options and details, see [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]].
=== Configure the repository as the configuration source ===
Add a <code>source_url</code> to <code>toolforge.yaml</code> that points to the copy of <code>toolforge.yaml</code> in your repository.
For example:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
config_version: v1beta1
source_url: https://gitlab.wikimedia.org/toolforge-repos/komla-test3/-/raw/main/toolforge.yaml
components:
recent-changes:
build:
repository: https://gitlab.wikimedia.org/toolforge-repos/komla-test3
ref: main
run:
command: python recent-changes.py
schedule: '@daily'
timeout: 3600
}}
The <code>source_url</code> tells Components Service to retrieve the tool configuration from the repository when deploying. This allows changes to <code>toolforge.yaml</code>, as well as changes to the application source code, to be managed through Git.
Commit and push the updated <code>toolforge.yaml</code> to the repository.
=== Create a deploy token ===
While logged in as your tool account, create a deploy token:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deploy-token create
}}
The command displays a token. Treat this token as a secret. Do not commit it to your repository.
=== Add the deploy token to GitLab ===
In your GitLab repository, go to:
<code>Settings → CI/CD → Variables</code>
Create a variable named:
{{Codesample|lang=text|scheme=light|code=
TOOLS_DEPLOY_TOKEN
}}
Set its value to the deploy token created in the previous step.
For security, make the variable hidden, protected, and masked where appropriate.
=== Add a GitLab CI configuration ===
Create a <code>.gitlab-ci.yml</code> file in the repository containing:
{{Codesample|lang=yaml|scheme=light|name=.gitlab-ci.yml|code=
include:
- project: "repos/cloud/cicd/gitlab-ci"
file: "toolforge-cd/deploy-to-toolforge.yaml"
}}
The included Toolforge CI configuration handles triggering the Components deployment.
This minimal configuration assumes that the GitLab project name and Toolforge tool name are the same. If they are different, or if you need to customize the deployment process, see [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]].
Commit and push <code>.gitlab-ci.yml</code> to the repository.
When a commit is pushed or merged to the default branch, GitLab CI will use the deploy token to request a new Components deployment for <code>komla-test3</code>.
When changes are pushed to the repository, GitLab CI will use the deploy token to trigger a Components deployment.
At this point, Push-to-Deploy is configured.
At this point, Push-to-Deploy is configured.
The setup gives the CI runner the credentials it needs to request a Components deployment for your tool.
After the CI configuration has been added to the repository, future deployments can follow this workflow:
{{Codesample|lang=text|scheme=light|code=
Git push
↓
CI runner
↓
Components Service deployment
↓
Toolforge workload
}}
You no longer need to SSH to Toolforge and manually run <code>toolforge components deployment create</code> for normal deployments.
== Test Push-to-Deploy ==
After configuring your CI runner, make a small change to <code>toolforge.yaml</code>.
For example, change the schedule:
{{Codesample|lang=yaml|scheme=light|name=toolforge.yaml|code=
run:
timeout: 1200
}}
Commit and push the change:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add toolforge.yaml
$ git commit -m "Change recent-changes schedule"
$ git push
}}
Do not manually create a deployment.
The push should trigger your CI runner, which creates a Components Service deployment automatically.
Check your repository's CI pipeline and verify that it completes successfully.
You can then inspect the deployment from Toolforge:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
For the scheduled example, verify that the new schedule has been applied:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
}}
If the deployment was triggered automatically and the new configuration was applied, Push-to-Deploy is working.
== Avoid independently managing the migrated job ==
After migrating a scheduled or continuous job, do not maintain a second independently managed copy of the same workload.
The component created by the Components Service may appear when you inspect workloads using Jobs Framework commands. This does not mean that you should continue managing that workload separately with <code>toolforge jobs</code> commands.
After migration, make changes to the component through <code>toolforge.yaml</code> and your source repository.
== Make future changes through Git ==
Once Push-to-Deploy is configured, make changes to your source code or <code>toolforge.yaml</code>, commit them, and push them to the configured Git branch.
For example:
{{Codesample|lang=shell-session|scheme=light|code=
$ git add .
$ git commit -m "Update recent-changes"
$ git push
}}
The CI runner triggers the deployment. You do not normally need to manually rebuild the Build Service image, update the Jobs Framework job, or create a Components deployment from the Toolforge bastion.
This allows the Git repository to track the source code and, when you choose to keep them together, the Toolforge deployment configuration.
== What changed? ==
Before migration, the deployment workflow is approximately:
{{Codesample|lang=text|scheme=light|code=
Git repository
↓
Build Service
↓
Build Service image
↓
Jobs Framework job
}}
The maintainer manages the build and the job separately.
After migrating to the Components Service and configuring Push-to-Deploy, the workflow becomes:
{{Codesample|lang=text|scheme=light|code=
Git repository
├── source code
└── toolforge.yaml
↓
Git push
↓
CI runner
↓
Components Service deployment
├── build
└── run
↓
Toolforge workload
}}
The Components configuration describes how the workload should be built and run, while Push-to-Deploy allows changes pushed to Git to trigger deployment automatically.
== Troubleshooting ==
=== Check the generated configuration ===
If the generated configuration does not look correct, do not deploy it immediately.
Compare it with:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge jobs list -o long
$ toolforge build list
}}
Correct <code>toolforge.yaml</code> before creating the Components configuration.
=== Check the Components configuration ===
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components config show
}}
Verify that the repository, ref, component name, command, schedule, and other settings are correct.
=== Check deployment history ===
Run:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment list
}}
Use the deployment information to determine whether the build or deployment failed.
=== Check your CI pipeline ===
If pushing a change does not create a deployment, check your repository's CI pipeline.
Verify that:
* the CI job ran after the Git push;
* the CI job completed successfully;
* the credentials required to trigger a Toolforge deployment are configured correctly.
See [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]] for the current setup instructions.
=== Rebuild the source manually ===
For troubleshooting, you can ask the Components Service to rebuild the source even when a matching build already exists:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-build
}}
This creates a manual deployment and bypasses the normal Push-to-Deploy workflow.
This also creates a manual deployment and is normally only needed for troubleshooting or administrative purposes.
=== Restart an unchanged component manually ===
If you need to recreate a component even though its configuration has not changed:
{{Codesample|lang=shell-session|scheme=light|code=
$ toolforge components deployment create --force-run
}}
=== Component fails to generate config from a running job ===
Make sure the job is using a build service image and not just a standard Toolforge image.
{{:Help:Cloud Services communication}}
== See also ==
* [[Help:Toolforge/Deploy your tool|Deploy your tool]]
* [[Help:Toolforge/Deploy your tool#Triggering a deployment from your CI runner|Triggering a deployment from your CI runner]]
* [[Help:Toolforge/Building container images|Building Toolforge container images]]
* [[Help:Toolforge/Running jobs|Running Toolforge jobs]]
[[Category:How-to-guide|{{SUBPAGENAME}}]]
iyepmb59feyaz70kt99xduchwxjplvj