Wikitech
labswiki
https://wikitech.wikimedia.org/wiki/Main_Page
MediaWiki 1.47.0-wmf.14
first-letter
Media
Special
Talk
User
User talk
Wikitech
Wikitech talk
File
File talk
MediaWiki
MediaWiki talk
Template
Template talk
Help
Help talk
Category
Category talk
Obsolete
Obsolete talk
OfficeIT
OfficeIT talk
Tool
Tool talk
Nova Resource
Nova Resource Talk
Heira
Heira Talk
TimedText
TimedText talk
Module
Module talk
Server Admin Log
0
7919
2445246
2445242
2026-08-09T02:01:20Z
Stashbot
7414
mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
2445246
wikitext
text/x-wiki
== 2026-08-09 ==
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-08 ==
* 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s)
* 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry)
* 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s)
* 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry)
* 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s)
* 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry)
* 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s)
* 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry)
* 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-07 ==
* 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
* 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]]
* 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]]
* 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]]
* 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet
* 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning
* 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning
* 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]]
* 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning
* 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]]
* 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]]
* 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json
* 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning
* 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning
* 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet
* 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]]
* 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows
* 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]]
* 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet
* 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet
* 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]]
* 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space
* 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-06 ==
* 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]]
* 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]]
* 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s)
* 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment
* 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]]
* 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s)
* 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment
* 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]]
* 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s)
* 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment
* 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]]
* 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie
* 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s)
* 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]])
* 18:18 ladsgroup@deploy1003: Stopping before sync operations
* 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]])
* 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage
* 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage
* 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage
* 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage
* 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]])
* 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]])
* 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm
* 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]])
* 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]])
* 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]])
* 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]])
* 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage
* 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage
* 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage
* 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage
* 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]]
* 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]]
* 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm
* 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm
* 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]]
* 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]]
* 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]]
* 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm
* 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]]
* 13:58 swfrench@dns1004: END - running authdns-update
* 13:56 swfrench@dns1004: START - running authdns-update
* 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm
* 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' .
* 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' .
* 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm
* 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' .
* 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' .
* 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' .
* 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' .
* 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
* 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
* 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
* 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' .
* 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' .
* 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' .
* 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm
* 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update
* 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update
* 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance
* 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update
* 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update
* 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json
* 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json
* 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command
* 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning
* 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning
* 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning
* 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet
* 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet
* 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet
* 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json
* 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet
* 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]]
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' .
* 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' .
* 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' .
* 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' .
* 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' .
* 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]]
* 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning
* 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning
* 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-05 ==
* 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm
* 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage
* 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage
* 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm
* 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet
* 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet
* 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm
* 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage
* 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage
* 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm
* 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet
* 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet
* 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad
* 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad
* 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad
* 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad
* 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006
* 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006
* 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm
* 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas
* 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage
* 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage
* 20:24 cjming: end of UTC late backport window
* 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s)
* 20:18 cjming@deploy1003: cjming: Continuing with deployment
* 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]]
* 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm
* 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s)
* 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm
* 20:08 jforrester@deploy1003: jforrester: Continuing with deployment
* 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]]
* 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]]
* 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003
* 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie
* 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]]
* 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm
* 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage
* 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet
* 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage
* 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet
* 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie
* 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003"
* 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003"
* 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad
* 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]]
* 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s)
* 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab
* 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad
* 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad
* 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad
* 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet
* 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet
* 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet
* 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet
* 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad
* 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]])
* 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]])
* 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]])
* 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]])
* 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]])
* 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]])
* 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet
* 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet
* 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet
* 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s)
* 15:58 reedy@deploy1003: reedy: Continuing with deployment
* 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]]
* 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie
* 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage
* 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage
* 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159
* 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159
* 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]]
* 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie
* 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]]
* 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159
* 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003"
* 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003"
* 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]]
* 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159
* 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie
* 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet
* 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet
* 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet
* 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet
* 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet
* 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet
* 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm
* 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]]
* 14:39 swfrench@dns1004: END - running authdns-update
* 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 14:37 swfrench@dns1004: START - running authdns-update
* 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie
* 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie
* 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage
* 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046
* 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046
* 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie
* 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage
* 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage
* 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage
* 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad
* 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad
* 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad
* 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage
* 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm
* 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie
* 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157
* 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157
* 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157
* 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003"
* 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003"
* 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s)
* 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 13:35 reedy@deploy1003: reedy: Continuing with deployment
* 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]]
* 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157
* 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie
* 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet
* 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet
* 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet
* 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet
* 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet
* 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet
* 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie
* 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage
* 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage
* 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 12:24 topranks: update bgp confed settings in eqsin
* 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156
* 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156
* 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156
* 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003"
* 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003"
* 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie
* 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
* 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
* 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156
* 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie
* 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet
* 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet
* 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet
* 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046
* 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046
* 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie
* 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet
* 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet
* 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync
* 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync
* 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync
* 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync
* 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie
* 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json
* 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning
* 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99)
* 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning
* 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage
* 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage
* 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155
* 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155
* 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]]
* 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155
* 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003"
* 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003"
* 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning
* 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning
* 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155
* 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie
* 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet
* 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json
* 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json
* 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet
* 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet
* 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet
* 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet
* 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet
* 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json
* 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet
* 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]]
* 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8
* 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5
* 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]]
* 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7
* 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2
* 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]]
* 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1
* 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]]
* 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8
* 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5
* 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]]
* 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6
* 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4
* 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json
* 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]]
* 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7
* 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2
* 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]]
* 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]]
* 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1
* 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json
* 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance
* 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json
* 07:21 slyngshede@dns1004: END - running authdns-update
* 07:19 slyngshede@dns1004: START - running authdns-update
* 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json
* 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json
* 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json
* 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json
* 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance
* 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json
* 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json
* 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json
* 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json
* 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json
* 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance
* 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance
* 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance
* 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json
* 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json
* 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json
* 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json
* 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json
* 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance
* 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json
* 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json
* 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json
* 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json
* 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance
* 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance
* 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json
* 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json
* 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json
* 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json
== 2026-08-04 ==
* 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json
* 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance
* 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json
* 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json
* 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json
* 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json
* 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json
* 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance
* 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance
* 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s)
* 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment
* 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm
* 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]]
* 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance
* 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json
* 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002"
* 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002"
* 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json
* 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046
* 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm
* 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json
* 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json
* 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}}
* 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json
* 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance
* 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json
* 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json
* 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet
* 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet
* 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet
* 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s)
* 17:48 swfrench@deploy1003: swfrench: Continuing with deployment
* 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]]
* 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json
* 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json
* 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s)
* 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure
* 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image
* 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue
* 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet
* 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit
* 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet
* 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/
* 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003
* 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json
* 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance
* 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:25 mutante: restarting gerrit - dropped outdated RSA host key
* 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003
* 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json
* 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json
* 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance
* 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia
* 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia
* 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
* 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key
* 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]]
* 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie
* 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s)
* 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]]
* 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s)
* 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]]
* 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment
* 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment
* 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment
* 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage
* 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage
* 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
* 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
* 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
* 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
* 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]]
* 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync
* 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync
* 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154
* 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154
* 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s)
* 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154
* 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003"
* 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003"
* 14:30 otto@deploy1003: otto: Continuing with deployment
* 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154
* 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie
* 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet
* 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]]
* 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet
* 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet
* 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 14:13 swfrench@dns1004: END - running authdns-update
* 14:13 Msz2001: Finished deployments for UTC afternoon backport window
* 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s)
* 14:11 swfrench@dns1004: START - running authdns-update
* 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment
* 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser
* 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]]
* 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet
* 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet
* 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s)
* 13:45 swfrench@dns1004: END - running authdns-update
* 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment
* 13:43 swfrench@dns1004: START - running authdns-update
* 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]]
* 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet
* 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet
* 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet
* 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet
* 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet
* 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet
* 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet
* 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie
* 13:05 swfrench@dns1004: END - running authdns-update
* 13:03 swfrench@dns1004: START - running authdns-update
* 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom
* 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom
* 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning
* 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie
* 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie
* 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage
* 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie
* 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage
* 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet
* 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet
* 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet
* 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet
* 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet
* 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie
* 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet
* 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet
* 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage
* 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie
* 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage
* 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet
* 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie
* 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet
* 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet
* 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows
* 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad
* 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet
* 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet
* 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad
* 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141
* 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141
* 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141
* 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003"
* 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003"
* 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet
* 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet
* 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad
* 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage
* 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie
* 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet
* 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet
* 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage
* 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141
* 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie
* 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet
* 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet
* 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26
* 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet
* 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet
* 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet
* 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie
* 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet
* 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet
* 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet
* 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet
* 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie
* 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet
* 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet
* 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet
* 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet
* 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie
* 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie
* 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie
* 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie
* 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie
* 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie
* 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet
* 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie
* 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie
* 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage
* 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage
* 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie
* 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage
* 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage
* 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage
* 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage
* 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage
* 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage
* 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie
* 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage
* 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage
* 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140
* 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140
* 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140
* 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie
* 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003"
* 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003"
* 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071
* 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie
* 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070
* 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie
* 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139
* 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139
* 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie
* 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003"
* 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003"
* 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie
* 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie
* 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069
* 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140
* 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie
* 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet
* 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139
* 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet
* 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet
* 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie
* 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet
* 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet
* 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet
* 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie
* 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage
* 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage
* 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071
* 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070
* 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie
* 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069
* 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet
* 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie
* 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet
* 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet
* 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie
* 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet
* 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet
* 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie
* 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet
* 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet
* 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet
* 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet
* 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet
* 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet
* 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage
* 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet
* 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage
* 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage
* 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage
* 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet
* 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie
* 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet
* 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet
* 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet
* 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet
* 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie
* 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet
* 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet
* 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie
* 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet
* 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet
* 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet
* 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie
* 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet
* 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]]
* 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet
* 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage
* 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage
* 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage
* 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage
* 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]]
* 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie
* 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie
* 06:50 slyngshede@dns1004: END - running authdns-update
* 06:48 slyngshede@dns1004: START - running authdns-update
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s)
* 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s)
* 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s)
* 00:41 cjming@deploy1003: cjming: Continuing with deployment
* 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]]
== 2026-08-03 ==
* 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync
* 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync
* 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie
* 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage
* 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage
* 21:42 dancy@deploy1003: Stopping before sync operations
* 21:41 dancy@deploy1003: Started scap sync-world: testing
* 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts
* 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s)
* 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s)
* 21:33 dancy@deploy1003: dancy: Continuing with deployment
* 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]]
* 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie
* 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s)
* 20:59 dancy@deploy1003: dancy: Continuing with deployment
* 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]]
* 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s)
* 20:48 cjming@deploy1003: cjming: Continuing with deployment
* 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]]
* 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s)
* 20:38 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]]
* 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s)
* 20:12 krinkle@deploy1003: krinkle: Continuing with deployment
* 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]]
* 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw
* 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s)
* 18:53 krinkle@deploy1003: krinkle: Continuing with deployment
* 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw
* 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]]
* 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s)
* 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet
* 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie
* 18:33 krinkle@deploy1003: krinkle: Continuing with deployment
* 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]]
* 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage
* 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage
* 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie
* 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors
* 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors
* 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox
* 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet
* 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet
* 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie
* 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage
* 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage
* 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors
* 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie
* 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie
* 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox
* 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet
* 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage
* 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage
* 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage
* 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage
* 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie
* 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie
* 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie
* 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage
* 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage
* 15:33 jhathaway@dns1004: END - running authdns-update
* 15:31 jhathaway@dns1004: START - running authdns-update
* 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie
* 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts
* 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s)
* 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie
* 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json
* 15:09 dancy@deploy1003: Started scap sync-world: testing
* 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts
* 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage
* 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s)
* 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage
* 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie
* 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie
* 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage
* 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage
* 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie
* 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie
* 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie
* 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage
* 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s)
* 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage
* 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage
* 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment
* 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage
* 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]]
* 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie
* 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie
* 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet
* 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet
* 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet
* 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet
* 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet
* 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie
* 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie
* 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet
* 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet
* 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet
* 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s)
* 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage
* 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage
* 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage
* 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage
* 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update
* 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie
* 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie
* 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet
* 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet
* 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie
* 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie
* 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network
* 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 11:19 marostegui@dns1004: END - running authdns-update
* 11:17 marostegui@dns1004: START - running authdns-update
* 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage
* 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply
* 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply
* 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage
* 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply
* 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply
* 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply
* 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply
* 11:02 marostegui@dns1004: END - running authdns-update
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage
* 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage
* 11:00 marostegui@dns1004: START - running authdns-update
* 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s)
* 10:51 cmooney@dns3003: END - running authdns-update
* 10:49 cmooney@dns3003: START - running authdns-update
* 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie
* 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003"
* 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003"
* 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]]
* 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json
* 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json
* 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]])
* 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie
* 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie
* 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage
* 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage
* 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage
* 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage
* 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie
* 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie
* 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie
* 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie
* 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage
* 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage
* 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage
* 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage
* 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie
* 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie
* 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash
* 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie
* 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie
* 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage
* 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage
* 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]])
* 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage
* 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage
* 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s)
* 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie
* 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie
* 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash
* 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]]
* 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]]
* 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-02 ==
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-01 ==
* 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-31 ==
* 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing
* 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox
* 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing
* 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing
* 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing
* 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie
* 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing
* 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing
* 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing
* 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515
* 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515
* 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie
* 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage
* 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage
* 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie
* 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts
* 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts
* 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048
* 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048
* 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage
* 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage
* 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie
* 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s)
* 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment
* 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]]
== 2026-07-30 ==
* 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts
* 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s)
* 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts
* 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s)
* 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]]
* 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s)
* 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment
* 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]]
* 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service
* 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie
* 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage
* 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage
* 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie
* 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie
* 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]]
* 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage
* 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]]
* 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage
* 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie
* 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]]
* 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie
* 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance
* 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage
* 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json
* 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance
* 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance
* 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage
* 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie
* 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie
* 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie
* 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage
* 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage
* 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance
* 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie
* 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance
* 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance
* 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json
* 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance
* 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie
* 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance
* 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s)
* 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage
* 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage
* 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie
* 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie
* 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie
* 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003"
* 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie
* 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003"
* 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie
* 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie
* 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json
* 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance
* 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage
* 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage
* 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance
* 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie
* 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json
* 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance
* 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance
* 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s)
* 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment
* 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]]
* 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid
* 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid
* 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie
* 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage
* 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie
* 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage
* 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s)
* 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance
* 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance
* 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]]
* 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s)
* 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance
* 14:14 stran@deploy1003: stran: Continuing with deployment
* 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json
* 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance
* 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance
* 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie
* 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]]
* 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance
* 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance
* 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s)
* 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance
* 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json
* 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance
* 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance
* 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment
* 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]]
* 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance
* 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance
* 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance
* 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json
* 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance
* 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s)
* 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment
* 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance
* 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance
* 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie
* 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance
* 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003"
* 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003"
* 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]]
* 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance
* 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie
* 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json
* 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance
* 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage
* 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage
* 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s)
* 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]]
* 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance
* 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie
* 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json
* 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance
* 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance
* 12:05 ayounsi@dns1004: END - running authdns-update
* 12:02 ayounsi@dns1004: START - running authdns-update
* 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie
* 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie
* 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance
* 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage
* 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage
* 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage
* 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance
* 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance
* 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage
* 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance
* 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json
* 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance
* 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie
* 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie
* 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance
* 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json
* 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance
* 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance
* 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance
* 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie
* 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance
* 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json
* 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance
* 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance
* 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage
* 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox
* 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox
* 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary
* 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary
* 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance
* 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json
* 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance
* 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie
* 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie
* 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance
* 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts
* 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003
* 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003
* 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json
* 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance
* 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance
* 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance
* 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts
* 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance
* 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie
* 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance
* 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie
* 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance
* 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance
* 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance
* 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json
* 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance
* 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance
* 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye
* 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance
* 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json
* 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance
* 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage
* 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance
* 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage
* 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance
* 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json
* 07:26 klausman@dns2004: END - running authdns-update
* 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie
* 07:24 klausman@dns2004: START - running authdns-update
* 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet
* 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance
* 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json
* 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance
* 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance
* 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance
* 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json
* 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance
* 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance
* 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed
* 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json
* 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json
* 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance
* 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json
* 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json
* 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance
* 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance
* 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json
* 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json
* 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance
* 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json
* 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance
* 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox
* 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json
* 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]]
* 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]]
== 2026-07-29 ==
* 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue
* 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards
* 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]]
* 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]]
* 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie
* 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye
* 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]'
* 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards
* 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s)
* 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance
* 21:08 aaron@deploy1003: aaron: Continuing with deployment
* 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]]
* 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s)
* 20:48 aaron@deploy1003: aaron: Continuing with deployment
* 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]]
* 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance
* 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s)
* 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json
* 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance
* 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance
* 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment
* 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie
* 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]]
* 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye
* 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94
* 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s)
* 19:32 zabe@deploy1003: zabe: Continuing with deployment
* 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance
* 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]]
* 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json
* 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance
* 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance
* 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003"
* 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003"
* 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance
* 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json
* 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]])
* 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json
* 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm
* 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards
* 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json
* 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json
* 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance
* 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage
* 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json
* 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance
* 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance
* 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance
* 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage
* 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json
* 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet
* 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet
* 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance
* 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance
* 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen
* 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]])
* 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]]
* 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]])
* 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022
* 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022
* 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]]
* 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022
* 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003"
* 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003"
* 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]])
* 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]]
* 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022
* 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm
* 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]])
* 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards
* 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s)
* 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]]
* 17:47 swfrench@dns1004: END - running authdns-update
* 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json
* 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance
* 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance
* 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 17:45 swfrench@dns1004: START - running authdns-update
* 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance
* 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs
* 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json
* 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance
* 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance
* 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance
* 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance
* 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json
* 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance
* 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance
* 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s)
* 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674]
* 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s)
* 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s)
* 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674]
* 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s)
* 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674]
* 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true
* 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]]
* 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]]
* 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s)
* 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance
* 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards
* 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json
* 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance
* 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance
* 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance
* 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]]
* 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye
* 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance
* 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json
* 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance
* 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance
* 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet
* 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance
* 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment
* 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet
* 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json
* 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance
* 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance
* 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet
* 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json
* 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance
* 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie
* 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance
* 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance
* 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s)
* 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]]
* 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts
* 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674]
* 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm
* 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train
* 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage
* 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s)
* 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm
* 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage
* 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]]
* 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment
* 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance
* 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]]
* 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json
* 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance
* 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage
* 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance
* 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie
* 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance
* 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie
* 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage
* 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance
* 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json
* 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance
* 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance
* 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance
* 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage
* 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json
* 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance
* 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance
* 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance
* 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s)
* 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json
* 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance
* 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance
* 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards
* 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s)
* 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json
* 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance
* 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance
* 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm
* 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage
* 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm
* 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment
* 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie
* 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie
* 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet
* 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie
* 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021
* 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003"
* 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003"
* 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]]
* 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage
* 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance
* 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]]
* 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage
* 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]]
* 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json
* 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance
* 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance
* 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage
* 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org
* 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance
* 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye
* 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye
* 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance
* 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie
* 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]])
* 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org
* 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098']
* 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098']
* 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097']
* 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097']
* 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage
* 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json
* 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance
* 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance
* 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance
* 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors
* 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors
* 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors
* 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors
* 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json
* 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance
* 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance
* 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance
* 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021
* 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015
* 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015
* 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json
* 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance
* 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance
* 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm
* 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s)
* 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm
* 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]])
* 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json
* 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance
* 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance
* 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage
* 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage
* 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie
* 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie
* 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm
* 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285
* 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285
* 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie
* 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance
* 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002"
* 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002"
* 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json
* 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance
* 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox
* 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance
* 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage
* 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance
* 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad
* 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance
* 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage
* 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json
* 14:03 sukhe@dns1004: END - running authdns-update
* 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance
* 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance
* 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance
* 14:01 sukhe@dns1004: START - running authdns-update
* 14:00 sukhe@dns1004: START - running authdns-update
* 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance
* 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json
* 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance
* 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance
* 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader
* 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json
* 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance
* 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance
* 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie
* 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json
* 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance
* 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance
* 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply
* 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service
* 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply
* 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie
* 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm
* 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage
* 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s)
* 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage
* 13:35 stran@deploy1003: stran: Continuing with deployment
* 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t
* 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage
* 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet
* 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]]
* 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad
* 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage
* 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad
* 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage
* 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s)
* 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet
* 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet
* 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet
* 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage
* 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265
* 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265
* 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie
* 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance
* 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet
* 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment
* 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet
* 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet
* 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad
* 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]]
* 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json
* 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance
* 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet
* 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance
* 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s)
* 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance
* 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet
* 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet
* 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet
* 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance
* 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance
* 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment
* 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance
* 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]]
* 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json
* 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance
* 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003"
* 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance
* 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet
* 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json
* 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie
* 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance
* 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance
* 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003"
* 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie
* 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance
* 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json
* 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance
* 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance
* 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json
* 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance
* 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance
* 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]]
* 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet
* 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet
* 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet
* 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet
* 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]]
* 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet
* 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet
* 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet
* 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]]
* 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox
* 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet
* 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet
* 12:38 ayounsi@dns1004: END - running authdns-update
* 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet
* 12:35 ayounsi@dns1004: START - running authdns-update
* 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet
* 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet
* 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet
* 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance
* 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance
* 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json
* 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance
* 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance
* 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts
* 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance
* 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json
* 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance
* 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance
* 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance
* 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync
* 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync
* 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance
* 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]]
* 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance
* 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json
* 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance
* 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json
* 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance
* 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json
* 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance
* 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance
* 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json
* 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance
* 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts
* 12:00 marostegui: Rename tables [[phab:T425074|T425074]]
* 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance
* 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox
* 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync
* 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync
* 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet
* 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync
* 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync
* 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance
* 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance
* 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json
* 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance
* 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance
* 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json
* 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance
* 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance
* 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance
* 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json
* 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance
* 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance
* 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]]
* 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance
* 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance
* 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]])
* 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance
* 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json
* 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance
* 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance
* 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json
* 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance
* 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance
* 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json
* 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance
* 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance
* 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance
* 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json
* 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]]
* 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]]
* 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance
* 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet
* 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance
* 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance
* 09:44 btullis@dns1004: END - running authdns-update
* 09:42 btullis@dns1004: START - running authdns-update
* 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json
* 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance
* 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance
* 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet
* 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance
* 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json
* 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance
* 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance
* 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json
* 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance
* 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance
* 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance
* 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8
* 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5
* 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8
* 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5
* 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]]
* 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]]
* 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org
* 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade
* 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org
* 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org
* 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org
* 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org
* 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org
* 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade
* 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]]
* 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance
* 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance
* 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance
* 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance
* 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade
* 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie
* 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json
* 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance
* 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance
* 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json
* 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance
* 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance
* 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json
* 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance
* 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance
* 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json
* 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance
* 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance
* 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes
* 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage
* 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts
* 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie
* 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance
* 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance
* 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance
* 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance
* 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json
* 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance
* 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json
* 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json
* 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance
* 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance
* 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json
* 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance
* 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet
* 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet
* 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie
* 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage
* 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage
* 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie
* 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage
* 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie
* 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage
* 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage
* 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie
* 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage
== 2026-07-28 ==
* 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet
* 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet
* 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet
* 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards
* 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards
* 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s)
* 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards
* 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards
* 20:54 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s)
* 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s)
* 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]]
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s)
* 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:30 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]]
* 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm
* 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]]
* 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s)
* 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm
* 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]]
* 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance
* 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage
* 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage
* 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage
* 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage
* 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014
* 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014
* 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm
* 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006
* 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006
* 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013
* 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003"
* 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003"
* 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance
* 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json
* 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance
* 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance
* 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts
* 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s)
* 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging]
* 18:29 sukhe@dns1004: END - running authdns-update
* 18:27 sukhe@dns1004: START - running authdns-update
* 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging]
* 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance
* 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json
* 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance
* 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance
* 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098
* 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098
* 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097
* 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097
* 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002"
* 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002"
* 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie
* 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:41 sukhe@dns1004: END - running authdns-update
* 17:39 sukhe@dns1004: START - running authdns-update
* 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid
* 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid
* 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005
* 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005
* 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003"
* 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003"
* 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance
* 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json
* 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance
* 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance
* 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie
* 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005
* 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005
* 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage
* 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage
* 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet
* 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet
* 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage
* 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage
* 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet
* 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie
* 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet
* 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet
* 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet
* 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie]
* 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts
* 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance
* 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138
* 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138
* 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138
* 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003"
* 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003"
* 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet
* 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json
* 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance
* 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet
* 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet
* 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet
* 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance
* 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards
* 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet
* 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet
* 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet
* 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet
* 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet
* 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet
* 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm
* 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance
* 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance
* 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138
* 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie
* 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade
* 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json
* 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance
* 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance
* 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts
* 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013
* 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance
* 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage
* 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm
* 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage
* 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet
* 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003
* 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet
* 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet
* 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s)
* 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]]
* 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts
* 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s)
* 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]]
* 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards
* 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts
* 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet
* 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet
* 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet
* 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main
* 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment
* 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment
* 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment
* 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance
* 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment
* 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm
* 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json
* 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance
* 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance
* 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance
* 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance
* 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts
* 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts
* 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet
* 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json
* 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance
* 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance
* 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet
* 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance
* 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards
* 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance
* 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json
* 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance
* 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance
* 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s)
* 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm
* 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet
* 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet
* 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet
* 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance
* 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]]
* 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet
* 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json
* 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance
* 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance
* 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance
* 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]]
* 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade
* 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json
* 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance
* 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]]
* 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance
* 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance
* 13:45 sukhe: restart pybal on A:lvs-codfw
* 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage
* 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance
* 13:44 btullis@dns1004: END - running authdns-update
* 13:42 sukhe: restart pybal on lvs2014
* 13:42 btullis@dns1004: START - running authdns-update
* 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json
* 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance
* 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance
* 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json
* 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance
* 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance
* 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]]
* 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]]
* 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]]
* 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards
* 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]]
* 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012
* 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012
* 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]]
* 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s)
* 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]]
* 13:18 swfrench@dns1004: END - running authdns-update
* 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment
* 13:16 swfrench@dns1004: START - running authdns-update
* 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012
* 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003
* 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]]
* 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance
* 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance
* 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003"
* 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003"
* 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s)
* 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s)
* 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox
* 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance
* 13:06 esanders@deploy1003: esanders: Continuing with deployment
* 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance
* 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance
* 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]]
* 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json
* 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance
* 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance
* 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance
* 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance
* 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json
* 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance
* 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance
* 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw
* 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet
* 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet
* 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003"
* 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003"
* 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet
* 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json
* 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance
* 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance
* 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json
* 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance
* 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance
* 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet
* 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet
* 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet
* 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet
* 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox
* 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet
* 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet
* 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet
* 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet
* 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet
* 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet
* 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet
* 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin
* 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet
* 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin
* 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance
* 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance
* 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie
* 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet
* 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet
* 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet
* 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance
* 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet
* 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json
* 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance
* 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet
* 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet
* 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet
* 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance
* 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance
* 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance
* 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet
* 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage
* 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json
* 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance
* 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance
* 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json
* 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance
* 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance
* 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage
* 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet
* 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet
* 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet
* 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet
* 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet
* 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet
* 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137
* 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003"
* 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet
* 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet
* 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet
* 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet
* 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet
* 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance
* 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance
* 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance
* 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet
* 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance
* 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet
* 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet
* 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance
* 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json
* 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance
* 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json
* 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance
* 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json
* 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance
* 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003"
* 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet
* 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet
* 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet
* 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet
* 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet
* 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s)
* 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137
* 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet
* 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw
* 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie
* 10:28 jforrester@deploy1003: jforrester: Continuing with deployment
* 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]]
* 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet
* 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet
* 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet
* 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance
* 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker
* 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet
* 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet
* 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet
* 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance
* 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet
* 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet
* 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad
* 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad
* 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet
* 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet
* 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet
* 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw
* 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw
* 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw
* 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance
* 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw
* 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw
* 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw
* 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw
* 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw
* 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance
* 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw
* 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance
* 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet
* 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance
* 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]]
* 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade
* 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]]
* 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json
* 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance
* 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance
* 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet
* 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet
* 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet
* 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet
* 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]]
* 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet
* 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker
* 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]]
* 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]]
* 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance
* 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance
* 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance
* 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance
* 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance
* 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance
* 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json
* 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance
* 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance
* 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json
* 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance
* 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade
* 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance
* 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json
* 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance
* 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance
* 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance
* 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance
* 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json
* 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance
* 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json
* 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance
* 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance
* 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json
* 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance
* 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]]
* 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie
* 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003"
* 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003"
* 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage
* 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage
* 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie
* 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie
* 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie
* 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie
* 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin
* 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin
* 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003"
* 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003"
* 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox
* 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]])
* 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie
== 2026-07-27 ==
* 23:50 Amir1: mass deleting vp8 transcodes
* 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye
* 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]]
* 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm
* 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye
* 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]]
* 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]]
* 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage
* 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage
* 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]]
* 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012
* 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm
* 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022
* 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022
* 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm
* 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004
* 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
* 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
* 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards
* 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022
* 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm
* 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards
* 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]]
* 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance
* 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s)
* 20:11 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]]
* 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s)
* 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance
* 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s)
* 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json
* 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance
* 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance
* 19:23 krinkle@deploy1003: krinkle: Continuing with deployment
* 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]]
* 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance
* 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance
* 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply
* 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply
* 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm
* 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm
* 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance
* 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s)
* 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json
* 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance
* 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance
* 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]]
* 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance
* 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage
* 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json
* 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance
* 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance
* 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage
* 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage
* 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage
* 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance
* 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json
* 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance
* 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance
* 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011
* 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003"
* 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003"
* 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011
* 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021
* 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021
* 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm
* 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm
* 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance
* 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json
* 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance
* 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance
* 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main
* 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance
* 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s)
* 17:27 taavi@deploy1003: taavi: Continuing with deployment
* 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json
* 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance
* 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance
* 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]]
* 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance
* 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json
* 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance
* 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance
* 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance
* 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json
* 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance
* 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance
* 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance
* 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance
* 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json
* 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance
* 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance
* 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance
* 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json
* 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance
* 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance
* 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance
* 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance
* 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance
* 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json
* 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance
* 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance
* 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance
* 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance
* 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json
* 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance
* 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json
* 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance
* 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s)
* 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance
* 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance
* 15:18 zabe@deploy1003: zabe: Continuing with deployment
* 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]]
* 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s)
* 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm
* 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance
* 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance
* 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance
* 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance
* 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json
* 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance
* 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json
* 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance
* 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance
* 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm
* 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance
* 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage
* 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage
* 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003
* 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance
* 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010
* 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003"
* 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003"
* 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage
* 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003
* 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage
* 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org
* 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance
* 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance
* 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance
* 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts
* 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance
* 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts
* 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet
* 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet
* 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet
* 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance
* 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json
* 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance
* 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance
* 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json
* 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance
* 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020
* 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020
* 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json
* 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance
* 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json
* 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm
* 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance
* 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm
* 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts
* 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts
* 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance
* 13:27 Lucas_WMDE: UTC afternoon backport+config window doen
* 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s)
* 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment
* 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]]
* 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]]
* 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]]
* 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]]
* 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance
* 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json
* 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance
* 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance
* 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors
* 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors
* 12:30 sukhe@dns1004: END - running authdns-update
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s)
* 12:28 sukhe@dns1004: START - running authdns-update
* 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]]
* 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts
* 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts
* 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]]
* 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool
* 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts
* 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts
* 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance
* 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]]
* 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade
* 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s)
* 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json
* 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance
* 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]]
* 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]]
* 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing
* 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing
* 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing
* 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing
* 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]]
* 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool
* 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]]
* 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing
* 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing
* 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json
* 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json
* 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing
* 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing
* 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json
* 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance
* 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance
* 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008
* 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie
* 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away
* 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance
* 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage
* 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json
* 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance
* 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage
* 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance
* 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136
* 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136
* 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136
* 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003"
* 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003"
* 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136
* 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie
* 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]]
* 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet
* 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance
* 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet
* 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet
* 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json
* 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance
* 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance
* 07:44 phuedx: UTC morning backport window done
* 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s)
* 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance
* 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]]
* 07:25 phuedx@deploy1003: phuedx: Continuing with deployment
* 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]]
* 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json
* 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance
* 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete
* 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]]
* 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet
* 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet
* 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning
* 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting
* 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards
* 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards
* 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards
* 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards
* 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s)
* 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s)
* 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-26 ==
* 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm
* 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm
* 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage
* 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage
* 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage
* 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage
* 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019
* 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019
* 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018
* 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018
* 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm
* 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm
== 2026-07-25 ==
* 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering
* 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet
* 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet
* 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards
* 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards
* 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards
* 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards
* 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s)
* 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s)
* 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm
* 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm
* 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage
* 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage
* 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage
* 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017
* 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024
* 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017
* 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003"
* 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003"
* 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox
* 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox
* 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017
* 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024
* 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm
* 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm
* 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses
* 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet
* 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet
* 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards
* 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards
* 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards
* 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards
* 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s)
* 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s)
* 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm
* 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage
* 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage
* 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023
* 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023
* 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023
* 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003"
* 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003"
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm
* 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage
* 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage
* 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox
* 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016
* 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016
* 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023
* 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm
* 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm
* 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday
== 2026-07-24 ==
* 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet
* 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet
* 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie
* 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage
* 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage
* 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie
* 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie
* 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet
* 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet
* 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet
* 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards
* 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie
* 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts
* 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts
* 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad
* 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage
* 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad
* 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage
* 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm
* 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135
* 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135
* 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage
* 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage
* 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s)
* 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003"
* 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003"
* 14:29 krinkle@deploy1003: krinkle: Continuing with deployment
* 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135
* 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie
* 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet
* 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet
* 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet
* 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet
* 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet
* 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet
* 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet
* 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet
* 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet
* 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet
* 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet
* 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]]
* 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003"
* 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003"
* 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors
* 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors
* 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm
* 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:30 papaul: reboot mr1-eqsin for maintenance
* 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet
* 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet
* 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts
* 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts
* 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts
* 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts
* 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts
* 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts
* 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie
* 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts
* 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet
* 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet
* 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet
* 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet
* 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet
* 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet
* 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet
* 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet
* 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet
* 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet
* 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet
* 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet
* 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet
* 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet
* 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet
* 09:30 brouberol@dns1004: END - running authdns-update
* 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet
* 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet
* 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts
* 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet
* 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet
* 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts
* 09:26 brouberol@dns1004: START - running authdns-update
* 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts
* 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet
* 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet
* 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet
* 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts
* 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet
* 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts
* 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet
* 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet
* 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet
* 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet
* 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet
* 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet
* 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet
* 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning
* 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6
* 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6
* 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4
* 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue
* 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-23 ==
* 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh
* 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003
* 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003
* 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards
* 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s)
* 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003
* 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003
* 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s)
* 20:13 dani@deploy1003: dani: Continuing with deployment
* 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]]
* 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002"
* 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002"
* 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm
* 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99)
* 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop
* 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage
* 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage
* 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards
* 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s)
* 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014
* 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003"
* 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003"
* 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance
* 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox
* 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance
* 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance
* 17:29 cmooney@dns3003: END - running authdns-update
* 17:27 cmooney@dns3003: START - running authdns-update
* 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance
* 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance
* 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance
* 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json
* 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json
* 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]]
* 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json
* 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]]
* 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128
* 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128
* 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance
* 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance
* 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet
* 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie
* 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance
* 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage
* 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage
* 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072
* 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072
* 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie
* 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance
* 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s)
* 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance
* 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]]
* 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet
* 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet
* 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet
* 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003"
* 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003"
* 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 15:19 cmooney@dns2004: END - running authdns-update
* 15:17 cmooney@dns2004: START - running authdns-update
* 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org
* 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie
* 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet
* 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie
* 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm
* 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance
* 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org
* 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet
* 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet
* 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance
* 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet
* 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet
* 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled
* 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet
* 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage
* 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage
* 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage
* 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage
* 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]]
* 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage
* 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage
* 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance
* 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie
* 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet
* 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071
* 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071
* 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards
* 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance
* 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance
* 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance
* 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071
* 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003"
* 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003"
* 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance
* 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance
* 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet
* 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet
* 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance
* 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance
* 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance
* 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance
* 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014
* 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance
* 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm
* 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet
* 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008
* 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008
* 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org
* 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm
* 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]]
* 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade
* 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade
* 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie
* 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071
* 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie
* 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet
* 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet
* 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet
* 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]]
* 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance
* 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json
* 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance
* 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s)
* 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts
* 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance
* 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment
* 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]]
* 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s)
* 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies
* 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling
* 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance
* 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json
* 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance
* 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling
* 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling
* 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling
* 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie
* 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance
* 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet
* 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet
* 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet
* 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json
* 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance
* 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance
* 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375
* 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375
* 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage
* 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad
* 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad
* 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie
* 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading
* 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance
* 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing
* 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing
* 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing
* 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing
* 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing
* 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json
* 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance
* 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing
* 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie
* 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts
* 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage
* 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage
* 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing
* 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing
* 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing
* 10:02 marostegui@dns1004: END - running authdns-update
* 10:00 marostegui@dns1004: START - running authdns-update
* 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing
* 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing
* 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070
* 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070
* 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing
* 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing
* 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing
* 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070
* 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003"
* 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003"
* 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts
* 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070
* 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie
* 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet
* 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet
* 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet
* 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing
* 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing
* 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing
* 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing
* 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing
* 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing
* 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet
* 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet
* 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet
* 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie
* 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage
* 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage
* 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts
* 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069
* 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069
* 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s)
* 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069
* 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003"
* 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003"
* 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 07:24 jiji@deploy1003: jiji: Continuing with deployment
* 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules
* 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7
* 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2
* 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7
* 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2
* 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts
* 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts
* 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069
* 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie
* 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet
* 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts
* 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet
* 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet
* 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts
* 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address
* 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet
* 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet
* 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards
== 2026-07-22 ==
* 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003
* 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003
* 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot`
* 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s)
* 21:44 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]]
* 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s)
* 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert)
* 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment
* 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing
* 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]]
* 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm
* 19:40 mutante: gerrit - one more service restart is needed - restarting
* 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing
* 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing
* 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing
* 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage
* 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s)
* 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation
* 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage
* 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards
* 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016
* 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016
* 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm
* 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s)
* 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s)
* 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d]
* 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s)
* 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated)
* 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d]
* 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s)
* 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d]
* 17:56 Raine: deployment server switchover => deploy1003 is primary now
* 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s)
* 17:54 mutante: restarting gerrit for maintenance
* 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm
* 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]]
* 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s)
* 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards
* 17:04 Raine: point deployment.eqiad.wmnet to deploy1003
* 17:04 kamila@dns7001: END - running authdns-update
* 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage
* 17:02 kamila@dns7001: START - running authdns-update
* 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover
* 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage
* 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]]
* 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s)
* 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]]
* 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s)
* 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]]
* 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013
* 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013
* 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013
* 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003"
* 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003"
* 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox
* 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013
* 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm
* 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards
* 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply
* 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply
* 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply
* 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply
* 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s)
* 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply
* 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply
* 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply
* 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply
* 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm
* 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s)
* 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]]
* 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s)
* 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]]
* 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s)
* 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]]
* 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage
* 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet
* 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet
* 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet
* 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage
* 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm
* 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s)
* 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]]
* 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013
* 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]]
* 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet
* 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013
* 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014
* 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal'
* 14:24 sukhe: restart pybal on lvs1019
* 14:24 sukhe: restart pybal on lvs1020
* 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts
* 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s)
* 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment
* 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet
* 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]]
* 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet
* 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s)
* 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test
* 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment
* 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet
* 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024']
* 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]]
* 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024']
* 13:26 sukhe@dns1004: END - running authdns-update
* 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 13:24 sukhe@dns1004: START - running authdns-update
* 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s)
* 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 13:18 stran@deploy2003: stran: Continuing with deployment
* 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]]
* 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test
* 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test
* 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test
* 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts
* 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet
* 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003
* 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet
* 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie
* 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s)
* 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage
* 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage
* 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]]
* 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates
* 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates
* 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie
* 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068
* 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068
* 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068
* 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003"
* 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003"
* 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage
* 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068
* 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie
* 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage
* 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet
* 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet
* 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet
* 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s)
* 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment
* 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie
* 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet
* 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet
* 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]]
* 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates
* 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]]
* 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates
* 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates
* 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s)
* 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 10:50 zabe@deploy2003: zabe: Continuing with deployment
* 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]]
* 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates
* 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates
* 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s)
* 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie
* 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]]
* 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage
* 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage
* 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie
* 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates
* 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet
* 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet
* 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]]
* 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003"
* 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003
* 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003
* 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003"
* 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates
* 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates
* 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates
* 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates
* 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning
* 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]]
* 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie
* 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage
* 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage
* 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s)
* 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1
* 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie
* 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]]
* 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s)
* 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet
* 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet
* 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment
* 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]]
* 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]]
* 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade
* 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade
* 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade
* 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade
* 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade
* 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1
* 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade
* 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade
* 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1
* 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1
* 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:47 phuedx: End of UTC morning backport window
* 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s)
* 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:39 phuedx@deploy2003: phuedx: Continuing with deployment
* 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]]
* 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003"
* 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003
* 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003
* 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003"
* 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation
* 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet
* 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5
* 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-21 ==
* 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards
* 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive
* 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s)
* 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024']
* 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024']
* 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024']
* 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards
* 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s)
* 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm
* 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm
* 20:52 krinkle@deploy2003: krinkle: Continuing with deployment
* 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]]
* 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s)
* 20:44 krinkle@deploy2003: krinkle: Rolling back deployment
* 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]]
* 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet
* 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet
* 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s)
* 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet
* 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet
* 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage
* 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet
* 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet
* 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 20:29 dani@deploy2003: dani: Continuing with deployment
* 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage
* 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet
* 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet
* 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet
* 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet
* 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]]
* 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet
* 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet
* 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s)
* 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm
* 20:10 sbisson@deploy2003: sbisson: Continuing with deployment
* 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]]
* 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s)
* 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]])
* 19:59 zabe@deploy2003: zabe: Continuing with deployment
* 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]]
* 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s)
* 19:48 zabe@deploy2003: zabe: Continuing with deployment
* 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]]
* 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024']
* 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024']
* 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024']
* 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery
* 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation
* 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet
* 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet
* 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet
* 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003"
* 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003"
* 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox
* 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox
* 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm
* 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm
* 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s)
* 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates
* 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet
* 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet
* 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet
* 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet
* 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet
* 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet
* 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet
* 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet
* 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet
* 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet
* 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet
* 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates
* 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet
* 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet
* 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.*
* 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet
* 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie
* 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024
* 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024
* 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance
* 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage
* 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage
* 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage
* 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage
* 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance
* 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance
* 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm
* 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie
* 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm
* 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage
* 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm
* 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]]
* 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.*
* 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]]
* 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance
* 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance
* 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update
* 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update
* 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s)
* 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update
* 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad
* 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad
* 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance
* 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance
* 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache
* 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance
* 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet
* 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet
* 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet
* 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet
* 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts
* 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts
* 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s)
* 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade
* 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]]
* 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]]
* 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet
* 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet
* 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update
* 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance
* 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance
* 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet
* 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade
* 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json
* 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet
* 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet
* 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet
* 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet
* 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet
* 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki
* 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade
* 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards
* 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json
* 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage
* 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade
* 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage
* 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json
* 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet
* 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet
* 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]]
* 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet
* 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet
* 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet
* 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet
* 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json
* 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012
* 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012
* 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm
* 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json
* 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance
* 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json
* 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s)
* 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 13:18 kharlan@deploy2003: kharlan: Continuing with deployment
* 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json
* 13:17 brouberol@dns1004: END - running authdns-update
* 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:15 brouberol@dns1004: START - running authdns-update
* 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]]
* 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json
* 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards
* 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json
* 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json
* 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json
* 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json
* 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json
* 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance
* 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json
* 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]]
* 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json
* 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json
* 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts
* 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts
* 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts
* 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts
* 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json
* 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json
* 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm
* 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json
* 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance
* 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json
* 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json
* 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json
* 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json
* 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance
* 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json
* 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json
* 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json
* 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance
* 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json
* 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json
* 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json
* 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json
* 11:21 XioNoX: put eqiad-drmrs Arelion link in service
* 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json
* 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json
* 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts
* 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json
* 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance
* 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json
* 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json
* 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json
* 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance
* 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json
* 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json
* 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json
* 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json
* 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json
* 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json
* 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel
* 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json
* 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance
* 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance
* 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json
* 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json
* 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance
* 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel
* 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet
* 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet
* 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet
* 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet
* 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json
* 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json
* 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]]
* 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json
* 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]]
* 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1
* 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003
* 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003
* 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s)
* 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment
* 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts
* 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org
* 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]]
* 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org
* 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades
* 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades
* 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync
* 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync
* 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync
* 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync
* 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync
* 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync
* 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage
* 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage
* 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003"
* 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003
* 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003
* 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003"
* 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning
* 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1
* 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8
* 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5
* 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s)
* 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s)
* 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
== 2026-07-20 ==
* 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis
* 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s)
* 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment
* 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]]
* 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards
* 22:39 maryum: Deployed security fixes for several security bugs
* 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards
* 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]]
* 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet
* 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet
* 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet
* 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s)
* 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage
* 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw
* 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage
* 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service
* 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts
* 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw
* 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s)
* 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat
* 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007
* 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007
* 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007
* 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003"
* 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003"
* 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json
* 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s)
* 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet
* 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet
* 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet
* 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet
* 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet
* 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet
* 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment
* 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json
* 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]]
* 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox
* 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json
* 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007
* 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm
* 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json
* 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json
* 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance
* 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json
* 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json
* 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm
* 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json
* 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json
* 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json
* 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance
* 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json
* 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json
* 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json
* 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json
* 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json
* 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance
* 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json
* 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s)
* 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s)
* 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json
* 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json
* 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json
* 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw
* 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet
* 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw
* 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet
* 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json
* 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance
* 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance
* 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host)
* 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet
* 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020']
* 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet
* 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage
* 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet
* 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet
* 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json
* 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage
* 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm
* 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet
* 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet
* 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet
* 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json
* 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s)
* 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet
* 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet
* 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011
* 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011
* 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm
* 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json
* 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json
* 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s)
* 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json
* 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance
* 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json
* 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json
* 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet
* 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet
* 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json
* 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json
* 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json
* 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance
* 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json
* 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020
* 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020
* 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020
* 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json
* 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox
* 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002"
* 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002"
* 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet
* 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet
* 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json
* 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox
* 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020
* 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm
* 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json
* 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards
* 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json
* 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance
* 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json
* 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet
* 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json
* 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet
* 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json
* 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023
* 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023
* 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json
* 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s)
* 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json
* 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance
* 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards
* 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment
* 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s)
* 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s)
* 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards
* 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet
* 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet
* 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm
* 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified]
* 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified]
* 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified]
* 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified]
* 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]]
* 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage
* 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards
* 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage
* 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm
* 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet
* 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet
* 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019
* 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019
* 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet
* 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage
* 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie
* 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet
* 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet
* 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie
* 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage
* 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet
* 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet
* 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet
* 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet
* 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet
* 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019
* 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003"
* 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003"
* 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet
* 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage
* 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019
* 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet
* 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet
* 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm
* 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027
* 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027
* 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm
* 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage
* 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw
* 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw
* 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet
* 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage
* 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage
* 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet
* 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet
* 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts
* 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts
* 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts
* 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts
* 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie
* 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie
* 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts
* 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts
* 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts
* 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts
* 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie
* 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie
* 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts
* 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts
* 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet
* 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet
* 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage
* 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet
* 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage
* 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage
* 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage
* 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet
* 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie
* 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts
* 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts
* 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts
* 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage
* 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie
* 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage
* 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage
* 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage
* 10:06 blake@deploy2003: Stopping before sync operations
* 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values
* 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie
* 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts
* 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts
* 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage
* 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts
* 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage
* 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie
* 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie
* 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s)
* 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values
* 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage
* 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage
* 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie
* 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie
* 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie
* 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage
* 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage
* 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage
* 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage
* 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie
* 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie
* 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down
* 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-18 ==
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards
== 2026-07-17 ==
* 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards
* 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards
* 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards
* 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm
* 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm
* 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage
* 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage
* 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage
* 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018
* 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003"
* 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003"
* 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018
* 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm
* 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026
* 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026
* 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm
* 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards
* 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards
* 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards
* 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm
* 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards
* 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm
* 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s)
* 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s)
* 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s)
* 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage
* 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s)
* 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage
* 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s)
* 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage
* 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage
* 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm
* 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017
* 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017
* 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm
* 18:30 bking@dns1004: END - running authdns-update
* 18:28 bking@dns1004: START - running authdns-update
* 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet
* 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet
* 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet
* 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 17:46 dzahn@dns1006: END - running authdns-update
* 17:44 dzahn@dns1006: START - running authdns-update
* 17:44 dzahn@dns1006: END - running authdns-update
* 17:42 dzahn@dns1006: START - running authdns-update
* 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264
* 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264
* 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet
* 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet
* 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet
* 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s)
* 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment
* 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]]
* 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339
* 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339
* 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003"
* 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003"
* 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet
* 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet
* 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99)
* 13:24 blake@dns1004: END - running authdns-update
* 13:22 blake@dns1004: START - running authdns-update
* 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet
* 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet
* 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet
* 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts
* 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet
* 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet
* 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts
* 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet
* 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet
* 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet
* 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet
* 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet
* 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet
* 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie
* 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage
* 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage
* 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet
* 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338
* 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338
* 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338
* 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003"
* 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003"
* 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet
* 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338
* 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie
* 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet
* 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet
* 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet
* 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet
* 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet
* 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet
* 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie
* 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet
* 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage
* 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage
* 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet
* 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet
* 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337
* 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337
* 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003"
* 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003"
* 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337
* 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie
* 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet
* 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet
* 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet
* 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet
* 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet
* 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet
* 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet
* 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet
* 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet
* 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet
* 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.)
* 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.)
* 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie
* 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.)
* 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.)
* 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1
* 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn
* 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage
* 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage
* 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet
* 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet
* 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336
* 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336
* 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336
* 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003"
* 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003"
* 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336
* 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie
* 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet
* 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet
* 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet
* 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet
* 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet
* 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet
* 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet
* 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet
* 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie
* 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet
* 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet
* 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet
* 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet
* 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage
* 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335
* 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003"
* 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003"
* 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges
* 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335
* 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie
* 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet
* 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet
* 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet
* 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99)
* 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges
* 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet
* 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet
* 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet
* 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet
* 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet
* 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet
* 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster
* 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges
* 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet
* 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet
* 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards
* 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards
* 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply
* 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply
* 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply
* 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply
* 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply
* 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply
* 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards
* 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards
* 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet
* 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet
* 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet
* 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm
* 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm
* 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie
== 2026-07-16 ==
* 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards
* 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage
* 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage
* 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage
* 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage
* 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage
* 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage
* 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025
* 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025
* 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027
* 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274
* 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm
* 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003"
* 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003"
* 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm
* 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274
* 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie
* 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet
* 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet
* 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet
* 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there
* 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad
* 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet
* 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet
* 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet
* 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards
* 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie
* 22:56 Amir1: deleting echo notifications from 2015 in group0
* 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm
* 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage
* 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s)
* 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet
* 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet
* 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet
* 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage
* 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s)
* 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment
* 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]]
* 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage
* 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272
* 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272
* 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272
* 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003"
* 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003"
* 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage
* 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272
* 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie
* 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet
* 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet
* 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet
* 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet
* 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet
* 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet
* 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie
* 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s)
* 21:40 sbassett@deploy2003: sbassett: Continuing with deployment
* 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025
* 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025
* 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]]
* 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025
* 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003"
* 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003"
* 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s)
* 21:26 sbassett@deploy2003: sbassett: Continuing with deployment
* 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie
* 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage
* 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]]
* 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025
* 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm
* 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage
* 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage
* 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage
* 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271
* 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271
* 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271
* 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003"
* 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003"
* 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet
* 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet
* 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet
* 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271
* 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie
* 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet
* 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet
* 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet
* 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s)
* 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269
* 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269
* 20:36 aude@deploy2003: aude: Continuing with deployment
* 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]]
* 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down
* 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269
* 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003"
* 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003"
* 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons.
* 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons.
* 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json
* 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json
* 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]]
* 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet
* 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json
* 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]]
* 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet
* 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet
* 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet
* 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet
* 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet
* 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet
* 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet
* 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet
* 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet
* 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet
* 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards
* 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet
* 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet
* 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet
* 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269
* 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie
* 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet
* 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet
* 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet
* 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s)
* 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]]
* 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie
* 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet
* 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet
* 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet
* 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage
* 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards
* 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s)
* 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s)
* 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage
* 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment
* 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]]
* 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268
* 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268
* 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie
* 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s)
* 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268
* 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003"
* 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003"
* 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment
* 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268
* 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie
* 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]]
* 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage
* 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage
* 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet
* 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet
* 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet
* 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet
* 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet
* 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet
* 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003"
* 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003"
* 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s)
* 17:04 jiji@deploy2003: jiji: Continuing with deployment
* 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]]
* 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie
* 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267
* 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003"
* 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003"
* 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin
* 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet
* 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams
* 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet
* 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams
* 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet
* 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet
* 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet
* 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet
* 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage
* 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad
* 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet
* 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad
* 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet
* 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267
* 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage
* 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie
* 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet
* 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage
* 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet
* 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet
* 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage
* 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet
* 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6
* 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie
* 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin
* 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet
* 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet
* 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet
* 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s)
* 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie
* 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270
* 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270
* 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270
* 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003"
* 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003"
* 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet
* 16:07 jiji@deploy2003: jiji: Continuing with deployment
* 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet
* 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet
* 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet
* 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270
* 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie
* 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet
* 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]]
* 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet
* 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet
* 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet
* 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet
* 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage
* 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie
* 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage
* 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s)
* 15:37 jiji@deploy2003: jiji: Continuing with deployment
* 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6
* 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6
* 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet
* 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]]
* 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet
* 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet
* 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet
* 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage
* 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266
* 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266
* 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet
* 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet
* 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm
* 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266
* 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003"
* 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003"
* 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266
* 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264
* 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264
* 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie
* 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264
* 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003"
* 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003"
* 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad
* 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet
* 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet
* 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet
* 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet
* 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet
* 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264
* 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265
* 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265
* 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]]
* 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265
* 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003"
* 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003"
* 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet
* 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet
* 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage
* 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet
* 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet
* 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6
* 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet
* 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265
* 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie
* 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet
* 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet
* 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet
* 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet
* 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage
* 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet
* 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet
* 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet
* 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet
* 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet
* 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet
* 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet
* 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet
* 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet
* 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet
* 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet
* 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet
* 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons.
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet
* 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet
* 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet
* 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s)
* 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons.
* 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet
* 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet
* 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet
* 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet
* 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]]
* 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s)
* 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]]
* 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]]
* 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm
* 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade
* 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet
* 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet
* 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade
* 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet
* 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet
* 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6
* 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6
* 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6
* 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6
* 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6
* 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6
* 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6
* 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6
* 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet
* 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster
* 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet
* 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet
* 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet
* 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet
* 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet
* 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]]
* 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet
* 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet
* 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet
* 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet
* 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet
* 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet
* 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage
* 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet
* 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet
* 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage
* 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad
* 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox)
* 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org
* 13:27 sukhe@dns1004: END - running authdns-update
* 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet
* 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet
* 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet
* 13:25 sukhe@dns1004: START - running authdns-update
* 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet
* 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263
* 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263
* 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263
* 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet
* 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003"
* 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003"
* 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet
* 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org
* 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*
* 13:14 cdobbins@dns1004: END - running authdns-update
* 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s)
* 13:13 cdobbins@dns1004: START - running authdns-update
* 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update
* 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263
* 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie
* 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.*
* 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet
* 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet
* 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet
* 13:09 sbisson@deploy2003: sbisson: Continuing with deployment
* 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]]
* 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org
* 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet
* 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet
* 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet
* 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet
* 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet
* 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet
* 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org
* 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet
* 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet
* 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet
* 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet
* 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org
* 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org
* 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet
* 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet
* 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet
* 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org
* 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet
* 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet
* 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet
* 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet
* 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet
* 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet
* 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet
* 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org
* 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox)
* 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad
* 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad
* 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams
* 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams
* 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin
* 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin
* 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad
* 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad
* 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie
* 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet
* 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet
* 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie
* 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
* 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet
* 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage
* 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet
* 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage
* 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie
* 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage
* 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage
* 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage
* 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage
* 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet
* 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet
* 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie
* 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie
* 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie
* 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover
* 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie
* 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 08:57 tappof: bump space for prometheus k8s-dse in eqiad
* 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet
* 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet
* 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet
* 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet
* 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage
* 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage
* 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage
* 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover
* 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie
* 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie
* 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad
* 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie
* 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie
* 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover
* 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover
* 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json
* 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json
* 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]]
* 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json
* 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]]
* 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage
* 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad
* 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage
* 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org
* 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie
* 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie
* 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet
* 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards
* 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards
* 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards
* 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards
* 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s)
* 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage
* 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s)
* 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage
* 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
== 2026-07-15 ==
* 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm
* 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm
* 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage
* 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage
* 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie
* 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages
* 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm
* 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages
* 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet
* 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie
* 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002']
* 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm
* 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia
* 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons.
* 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons.
* 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad
* 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm
* 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun!
* 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001']
* 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001']
* 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001']
* 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001']
* 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm
* 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch
* 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet
* 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet
* 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie
* 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage
* 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage
* 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie
* 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262
* 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262
* 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262
* 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003"
* 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003"
* 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262
* 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet
* 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie
* 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet
* 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet
* 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet
* 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet
* 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]]
* 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad
* 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad
* 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo
* 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet
* 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs
* 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet
* 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo
* 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet
* 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs
* 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update
* 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]]
* 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
* 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet
* 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet
* 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet
* 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet
* 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw
* 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet
* 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet
* 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update
* 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet
* 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet
* 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet
* 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet
* 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:06 sukhe: sre.dns.roll-reboot to resume later
* 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox)
* 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org
* 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update
* 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet
* 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet
* 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet
* 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet
* 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet
* 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet
* 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw
* 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet
* 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet
* 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org
* 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet
* 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet
* 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet
* 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet
* 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org
* 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet
* 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet
* 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet
* 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet
* 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie
* 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet
* 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet
* 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet
* 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org
* 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet
* 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet
* 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet
* 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet
* 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie
* 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet
* 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org
* 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet
* 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage
* 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage
* 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org
* 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage
* 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage
* 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet
* 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw
* 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet
* 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org
* 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons.
* 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet
* 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie
* 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet
* 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons.
* 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes
* 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm
* 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie
* 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org
* 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad
* 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie
* 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet
* 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org
* 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad
* 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw
* 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet
* 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet
* 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet
* 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet
* 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet
* 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet
* 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet
* 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet
* 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw
* 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org
* 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet
* 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet
* 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet
* 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage
* 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]]
* 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet
* 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage
* 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org
* 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad
* 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie
* 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie
* 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org
* 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad
* 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]]
* 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm
* 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet
* 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet
* 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org
* 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet
* 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet
* 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install
* 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad
* 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes
* 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]]
* 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org
* 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad
* 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org
* 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet
* 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet
* 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet
* 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001
* 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001
* 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet
* 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org
* 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install
* 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet
* 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001
* 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003"
* 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003"
* 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet
* 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs
* 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs
* 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install
* 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad
* 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad
* 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org
* 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox)
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo
* 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo
* 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards
* 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001
* 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie
* 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards
* 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s)
* 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes
* 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage
* 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage
* 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw
* 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009
* 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003"
* 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003"
* 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009
* 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie
* 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet
* 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet
* 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet
* 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet
* 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet
* 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet
* 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes
* 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s)
* 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie
* 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment
* 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates
* 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates
* 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]]
* 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage
* 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage
* 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates
* 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates
* 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie
* 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet
* 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad
* 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad
* 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet
* 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010
* 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003"
* 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003"
* 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010
* 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage
* 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie
* 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet
* 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet
* 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet
* 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage
* 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet
* 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates
* 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates
* 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet
* 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet
* 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s)
* 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet
* 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet
* 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie
* 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie
* 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet
* 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet
* 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet
* 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet
* 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet
* 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet
* 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet
* 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage
* 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates
* 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates
* 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet
* 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage
* 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet
* 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet
* 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet
* 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet
* 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie
* 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]]
* 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet
* 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]]
* 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet
* 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011
* 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011
* 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003"
* 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003"
* 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet
* 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet
* 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage
* 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011
* 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie
* 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet
* 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet
* 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet
* 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage
* 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet
* 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates
* 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates
* 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet
* 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet
* 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet
* 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw
* 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad
* 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad
* 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie
* 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie
* 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie
* 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates
* 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates
* 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage
* 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage
* 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage
* 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage
* 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie
* 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates
* 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012
* 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates
* 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003"
* 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003"
* 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie
* 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012
* 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie
* 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet
* 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet
* 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet
* 08:30 elukey@dns1004: END - running authdns-update
* 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage
* 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates
* 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates
* 08:27 elukey@dns1004: START - running authdns-update
* 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet
* 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage
* 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet
* 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org
* 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie
* 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org
* 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org
* 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates
* 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie
* 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org
* 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org
* 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org
* 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie
* 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage
* 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates
* 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates
* 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org
* 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage
* 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
* 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage
* 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage
* 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013
* 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013
* 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates
* 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates
* 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013
* 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003"
* 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003"
* 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie
* 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013
* 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie
* 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s)
* 07:09 kharlan@deploy2003: kharlan: Continuing with deployment
* 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]]
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
== 2026-07-14 ==
* 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru
* 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet
* 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru
* 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet
* 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet
* 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet
* 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet
* 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet
* 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot
* 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance
* 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot
* 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet
* 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet
* 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie
* 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s)
* 20:24 sbassett@deploy2003: sbassett: Continuing with deployment
* 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage
* 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]]
* 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage
* 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s)
* 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment
* 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet
* 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]]
* 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie
* 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet
* 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet
* 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet
* 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s)
* 19:07 jforrester@deploy2003: jforrester: Continuing with deployment
* 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]]
* 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet
* 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator)
* 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s)
* 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet
* 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet
* 17:32 swfrench@deploy2003: swfrench: Continuing with deployment
* 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance
* 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance
* 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance
* 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image
* 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet
* 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia
* 16:50 sukhe: pool cp2046
* 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet
* 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent"
* 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot
* 16:32 mutante: contint1003 - main CI server - rebooting
* 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance
* 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance
* 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet
* 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet
* 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe
* 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet
* 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe
* 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie
* 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014
* 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003"
* 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003"
* 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003"
* 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003
* 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003
* 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003"
* 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014
* 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie
* 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie
* 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts
* 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance
* 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance
* 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s)
* 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie
* 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]])
* 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s)
* 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie
* 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment
* 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]]
* 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage
* 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage
* 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage
* 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage
* 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]]
* 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum
* 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie
* 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie
* 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance
* 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie
* 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie
* 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw
* 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync
* 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet
* 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet
* 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync
* 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org
* 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync
* 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync
* 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org
* 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org
* 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet
* 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet
* 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw
* 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw
* 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org
* 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage
* 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir
* 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync
* 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync
* 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage
* 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage
* 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage
* 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy
* 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough
* 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy
* 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum
* 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet
* 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox)
* 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org
* 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie
* 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]]
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie
* 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet
* 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet
* 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie
* 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet
* 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet
* 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org
* 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet
* 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet
* 13:18 elukey@dns1004: END - running authdns-update
* 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet
* 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet
* 13:16 elukey@dns1004: START - running authdns-update
* 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage
* 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8
* 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5
* 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5
* 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8
* 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet
* 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet
* 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet
* 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet
* 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage
* 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org
* 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet
* 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet
* 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet
* 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet
* 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org
* 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage
* 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet
* 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance
* 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru
* 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance
* 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru
* 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance
* 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org
* 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance
* 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance
* 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance
* 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet
* 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet
* 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw
* 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance
* 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage
* 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw
* 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw
* 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance
* 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade
* 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw
* 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw
* 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie
* 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie
* 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]]
* 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org
* 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox)
* 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet
* 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet
* 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum
* 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy
* 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet
* 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet
* 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet
* 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy
* 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir
* 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough
* 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie
* 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie
* 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org
* 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie
* 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet
* 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet
* 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet
* 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet
* 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet
* 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org
* 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org
* 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet
* 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage
* 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie
* 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org
* 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org
* 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet
* 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet
* 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet
* 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage
* 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs
* 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage
* 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org
* 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org
* 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]])
* 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet
* 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org
* 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535
* 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet
* 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet
* 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet
* 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage
* 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage
* 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage
* 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet
* 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet
* 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet
* 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet
* 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535
* 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage
* 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet
* 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet
* 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet
* 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet
* 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage
* 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535
* 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet
* 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet
* 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet
* 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet
* 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage
* 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003"
* 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003"
* 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet
* 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet
* 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet
* 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet
* 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067
* 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003"
* 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535
* 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie
* 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet
* 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie
* 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie
* 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync
* 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie
* 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync
* 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535
* 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie
* 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie
* 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie
* 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync
* 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync
* 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage
* 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s)
* 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage
* 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment
* 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage
* 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage
* 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]]
* 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet
* 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet
* 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie
* 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie
* 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie
* 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003"
* 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s)
* 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie
* 11:03 kharlan@deploy2003: kharlan: Continuing with deployment
* 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox
* 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]]
* 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s)
* 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage
* 10:52 marostegui@dns1004: START - running authdns-update
* 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage
* 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage
* 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage
* 10:43 kharlan@deploy2003: kharlan: Continuing with deployment
* 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie
* 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover
* 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067
* 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie
* 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie
* 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet
* 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet
* 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie
* 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet
* 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet
* 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet
* 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie
* 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]]
* 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie
* 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage
* 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage
* 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert
* 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage
* 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage
* 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129
* 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage
* 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage
* 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie
* 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie
* 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover
* 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie
* 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie
* 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet
* 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055
* 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055
* 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage
* 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage
* 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage
* 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage
* 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet
* 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet
* 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet
* 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s)
* 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8]
* 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie
* 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie
* 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s)
* 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie
* 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie
* 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8]
* 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s)
* 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055
* 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003"
* 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003"
* 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8]
* 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json
* 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox
* 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055
* 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie
* 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet
* 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet
* 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet
* 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json
* 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]]
* 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage
* 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json
* 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]]
* 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage
* 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage
* 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:29 marostegui@dns1004: END - running authdns-update
* 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet
* 08:27 marostegui@dns1004: START - running authdns-update
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet
* 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet
* 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet
* 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie
* 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot
* 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet
* 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet
* 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet
* 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet
* 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet
* 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet
* 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie
* 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki)
* 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet
* 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet
* 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet
* 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet
* 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet
* 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet
* 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet
* 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet
* 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet
* 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet
* 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet
* 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet
* 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors
* 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors
* 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet
* 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet
* 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet
* 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet
* 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage
* 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors
* 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet
* 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors
* 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage
* 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki)
* 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage
* 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage
* 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet
* 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw)
* 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet
* 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet
* 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet
* 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors
* 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors
* 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors
* 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors
* 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw)
* 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie
* 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie
* 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org
* 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org
* 06:25 marostegui@dns1004: END - running authdns-update
* 06:23 marostegui@dns1004: START - running authdns-update
* 06:22 marostegui@dns1004: END - running authdns-update
* 06:20 marostegui@dns1004: START - running authdns-update
* 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot
* 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync
* 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync
* 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
* 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
* 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
* 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
* 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:26 marostegui@dns1004: END - running authdns-update
* 05:24 marostegui@dns1004: START - running authdns-update
* 05:24 marostegui@dns1004: START - running authdns-update
* 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org
* 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org
* 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s)
* 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s)
* 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-13 ==
* 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet
* 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet
* 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs
* 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet
* 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet
* 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs
* 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]]
* 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]]
* 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]]
* 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s)
* 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]]
* 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts
* 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s)
* 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s)
* 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment
* 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there
* 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]]
* 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts
* 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie
* 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet
* 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet
* 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs
* 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet
* 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet
* 17:06 dzahn@dns1006: END - running authdns-update
* 17:04 dzahn@dns1006: START - running authdns-update
* 17:01 dzahn@dns1006: END - running authdns-update
* 16:59 dzahn@dns1006: START - running authdns-update
* 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s)
* 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]]
* 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]]
* 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]])
* 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test
* 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test
* 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test
* 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test
* 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts
* 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs
* 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s)
* 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync
* 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync
* 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync
* 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync
* 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync
* 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync
* 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
* 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
* 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
* 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
* 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117
* 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117
* 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet
* 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet
* 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet
* 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet
* 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet
* 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet
* 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet
* 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet
* 15:36 sukhe: restart pybal on lvs1020
* 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet
* 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet
* 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 14:28 marostegui@dns1004: END - running authdns-update
* 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:27 marostegui@dns1004: START - running authdns-update
* 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot
* 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie
* 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]]
* 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]]
* 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
* 14:02 marostegui@dns1004: END - running authdns-update
* 14:00 marostegui@dns1004: START - running authdns-update
* 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet
* 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet
* 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage
* 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.*
* 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage
* 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie
* 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs
* 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s)
* 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment
* 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers
* 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]]
* 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie
* 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s)
* 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]]
* 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage
* 12:23 Msz2001: Deployed changes to private code for Suggested Investigations
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage
* 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s)
* 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment
* 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]]
* 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s)
* 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]]
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie
* 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s)
* 11:55 zabe@deploy2003: zabe: Continuing with deployment
* 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]]
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie
* 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]]
* 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8
* 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3
* 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5
* 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage
* 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage
* 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023
* 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023
* 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023
* 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023
* 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json
* 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json
* 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie
* 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie
* 09:42 marostegui@dns1004: END - running authdns-update
* 09:40 marostegui@dns1004: START - running authdns-update
* 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage
* 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot
* 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie
* 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet
* 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet
* 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3
* 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet
* 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet
* 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing
* 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning
* 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5
* 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8
* 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet
* 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet
* 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet
* 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet
* 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage
* 08:07 marostegui@dns1004: END - running authdns-update
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet
* 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet
* 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:05 marostegui@dns1004: START - running authdns-update
* 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage
* 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet
* 08:05 marostegui@dns1004: START - running authdns-update
* 08:05 marostegui@dns1004: START - running authdns-update
* 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:04 marostegui@dns1004: START - running authdns-update
* 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org
* 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet
* 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet
* 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet
* 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet
* 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org
* 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:52 Msz2001: UTC morning backport+config window done
* 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie
* {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}}
* 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet
* 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment
* {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}}
* 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing
* {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}}
* 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s)
* 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org
* 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment
* 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org
* 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet
* 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet
* 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org
* 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]]
* 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org
* 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org
* 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org
* 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org
* 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org
* 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]]
* 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot
* 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot
* 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot
* 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot
* 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot
* 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning
* 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]]
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-12 ==
* 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json
* 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json
* 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]]
* 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json
* 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]]
* 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-11 ==
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-10 ==
* 19:12 jhathaway@dns1004: END - running authdns-update
* 19:10 jhathaway@dns1004: START - running authdns-update
* 18:23 mutante: vrts2002 rebooting (not the active host)
* 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts)
* 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw
* 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw
* 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]])
* 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet
* 17:08 jhathaway@dns1004: END - running authdns-update
* 17:07 jhathaway@dns1004: START - running authdns-update
* 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown
* 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet
* 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet
* 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet
* 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet
* 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet
* 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet
* 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet
* 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet
* 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet
* 16:25 mutante: gitlab-runners (production) rebooting cluster one by one
* 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting
* 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting
* 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance
* 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet
* 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet
* 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet
* 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet
* 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet
* 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet
* 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet
* 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet
* 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet
* 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet
* 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet
* 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet
* 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet
* 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet
* 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet
* 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet
* 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet
* 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet
* 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet
* 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet
* 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet
* 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet
* 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet
* 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet
* 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet
* 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet
* 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie
* 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet
* 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet
* 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet
* 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet
* 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet
* 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet
* 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet
* 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet
* 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet
* 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage
* 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet
* 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet
* 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet
* 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet
* 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet
* 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org
* 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054
* 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003"
* 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003"
* 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org
* 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox
* 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054
* 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie
* 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet
* 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet
* 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet
* 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie
* 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet
* 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet
* 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet
* 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet
* 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet
* 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet
* 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet
* 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw)
* 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade
* 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie
* 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts
* 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts
* 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie
* 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie
* 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet
* 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet
* 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw)
* 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw
* 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw
* 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm
* 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm
* 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet
* 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet
* 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw
* 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage
* 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]]
* 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage
* 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053
* 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053
* 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053
* 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003"
* 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003"
* 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox
* 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053
* 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie
* 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet
* 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet
* 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet
* 09:04 brouberol@dns1004: END - running authdns-update
* 09:03 brouberol@dns1004: START - running authdns-update
* 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s)
* 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea]
* 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s)
* 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea]
* 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s)
* 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea]
* 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}}
* 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet
* 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet
* 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade
* 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts
* 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts
* 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-09 ==
* 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s)
* 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment
* 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]]
* 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet
* 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet
* 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet
* 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie
* 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage
* 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage
* 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165
* 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002"
* 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002"
* 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox
* 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165
* 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie
* 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet
* 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet
* 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet
* 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:42 maryum: Deploy fix for [[phab:T431684|T431684]]
* 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s)
* 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002
* 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002
* 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie
* 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s)
* 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment
* 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm
* 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0)
* 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots
* 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003
* 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]]
* 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99)
* 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003
* 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99)
* 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet
* 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie
* 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0)
* 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet
* 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm
* 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet
* 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet
* 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage
* 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage
* 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99)
* 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie
* 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org
* 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org
* 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:09 mutante: zuul[12]00[123] - rebooting for maintenance
* 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster
* 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99)
* 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster
* 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster
* 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance
* 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster
* 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 16:49 mutante: planet1003/planet2003 - rebooting
* 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating
* 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 15:51 jynus: restarting backupmon1001
* 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts
* 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts
* 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts
* 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts
* 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade
* 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet
* 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet
* 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet
* 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet
* 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync
* 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync
* 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie
* 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade
* 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]]
* 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002"
* 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002"
* 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet
* 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet
* 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage
* 13:57 moritzm: installing requests security updates
* 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet
* 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet
* 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage
* 13:50 moritzm: installing python-cryptography security updates
* 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet
* 13:44 Msz2001: UTC afternoon config+backport window is done
* 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297
* 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet
* 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org
* 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s)
* 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org
* 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment
* 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052
* 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052
* 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052
* 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003"
* 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003"
* 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]]
* 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox
* 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052
* 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie
* 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet
* 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet
* 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet
* 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s)
* 13:13 jforrester@deploy2003: jforrester: Continuing with deployment
* 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]]
* 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org
* 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org
* 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org
* 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org
* 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004
* 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org
* 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade
* 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet
* 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet
* 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet
* 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet
* 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade
* 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet
* 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org
* 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet
* 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org
* 11:55 jmm@dns1004: END - running authdns-update
* 11:53 jmm@dns1004: START - running authdns-update
* 11:50 jmm@dns1004: END - running authdns-update
* 11:48 jmm@dns1004: START - running authdns-update
* 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org
* 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org
* 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet
* 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet
* 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet
* 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet
* 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org
* 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org
* 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org
* 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org
* 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org
* 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet
* 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org
* 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org
* 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org
* 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet
* 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003
* 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003
* 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet
* 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org
* 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org
* 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm
* 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org
* 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org
* 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org
* 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org
* 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org
* 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org
* 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet
* 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]]
* 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B
* 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]]
* 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage
* 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org
* 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003
* 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003
* 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B
* 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org
* 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org
* 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org
* 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage
* 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet
* 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org
* 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004
* 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org
* 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003"
* 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003"
* 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet
* 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet
* 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet
* 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox
* 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org
* 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004
* 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm
* 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm
* 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet
* 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org
* 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet
* 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet
* 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet
* 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet
* 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet
* 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet
* 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet
* 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet
* 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet
* 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet
* 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet
* 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet
* 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet
* 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet
* 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s)
* 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance
* 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet
* 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance
* 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment
* 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot
* 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]]
* 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003
* 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003
* 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003
* 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003"
* 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003"
* 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox
* 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003
* 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm
* 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm
* 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance
* 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance
* 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:31 hashar@deploy2003: Rolling back deployment
* 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]]
* 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm
* 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet
* 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet
* 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]]
* 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance
* 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4
* 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet
* 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet
* 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance
* 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance
* 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet
* 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance
* 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance
* 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance
* 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet
* 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4
* 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie
* 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s)
* 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment
* 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]]
* 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage
* 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage
* 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie
* 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie
* 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie
* 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie
* 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]]
* 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.*
* 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration
* 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad
* 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad
* 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-08 ==
* 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]])
* 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003"
* 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003"
* 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s)
* 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment
* 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]]
* 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie
* 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]])
* 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw
* 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw
* 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw
* 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw
* 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002
* 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003"
* 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003"
* 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002
* 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage
* 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage
* 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie
* 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye
* 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002"
* 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002"
* 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage
* 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage
* 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie
* 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]])
* 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s)
* 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye
* 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage
* 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage
* 20:30 cjming@deploy2003: cjming: Continuing with deployment
* 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie
* 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie
* 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]]
* 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage
* 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage
* 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw
* 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw
* 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw
* 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw
* 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003)
* 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet
* 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet
* 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003)
* 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002)
* 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc
* 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc
* 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie
* 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw
* 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw
* 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw
* 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw
* 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw
* 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw
* 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster
* 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie
* 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw
* 19:19 mutante: gerrit - replacing private key for registerEmail verification
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw
* 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw
* 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw
* 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw
* 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru
* 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw
* 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw
* 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw
* 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw
* 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw
* 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw
* 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw
* 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru
* 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru
* 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru
* 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru
* 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage
* 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie
* 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]]
* 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie
* 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie
* 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie
* 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage
* 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage
* 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage
* 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122
* 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122
* 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122
* 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003"
* 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003"
* 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage
* 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox
* 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122
* 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie
* 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie
* 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s)
* 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]]
* 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s)
* 17:42 kamila@deploy2003: Rolling back deployment
* 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s)
* 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 17:11 kamila@dns1005: END - running authdns-update
* 17:09 kamila@dns1005: START - running authdns-update
* 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet
* 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet
* 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet
* 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover
* 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server
* 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s)
* 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie
* 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server
* 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie
* 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie
* 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie
* 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage
* 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage
* 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage
* 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage
* 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage
* 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage
* 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test
* 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test
* 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test
* 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test
* 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164
* 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164
* 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164
* 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002"
* 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002"
* 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie
* 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox
* 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164
* 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie
* 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet
* 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet
* 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet
* 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie
* 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet
* 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet
* 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet
* 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet
* 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie
* 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage
* 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet
* 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage
* 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test
* 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet
* 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test
* 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test
* 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet
* 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet
* 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]]
* 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]]
* 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet
* 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet
* 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie
* 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]]
* 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]]
* 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service"
* 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie
* 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004
* 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004
* 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]]
* 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet
* 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet
* 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie
* 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test
* 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:29 swfrench@dns1004: END - running authdns-update
* 14:29 moritzm: installing jackson-core security updates
* 14:27 swfrench@dns1004: START - running authdns-update
* 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie
* 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 14:19 moritzm: installing librabbitmq security updates
* 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet
* 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet
* 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test
* 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet
* 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet
* 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie
* 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet
* 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad
* 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage
* 13:54 moritzm: installing libcap2 security updates
* 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage
* 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet
* 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0)
* 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool
* 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad
* 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet
* 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet
* 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet
* 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie
* 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119
* 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119
* 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003"
* 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003"
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage
* 13:39 moritzm: installing krb5 security updates
* 13:37 Lucas_WMDE: UTC afternoon backport+config window done
* 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie
* 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage
* 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox
* 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s)
* 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119
* 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie
* 13:30 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:30 moritzm: installing openssh security updates
* 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough
* 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]]
* 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie
* 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage
* 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118
* 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s)
* 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage
* 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage
* 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118
* 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003"
* 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003"
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough
* 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox
* 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118
* 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage
* 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]]
* 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw
* 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie
* 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage
* 13:04 moritzm: installing jq security updates
* 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage
* 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw
* 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051
* 12:43 moritzm: installing Python 3.11 security updates
* 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003"
* 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003"
* 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox
* 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051
* 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie
* 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet
* 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet
* 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet
* 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie
* 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie
* 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie
* 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet
* 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet
* 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]]
* 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage
* 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers
* 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage
* 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage
* 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage
* 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie
* 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie
* 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie
* 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie
* 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet
* 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage
* 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage
* 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage
* 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie
* 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet
* 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet
* 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie
* 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply
* 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply
* 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test
* 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance
* 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0)
* 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit
* 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie
* 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart
* 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99)
* 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart
* 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance
* 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet
* 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet
* 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet
* 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet
* 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet
* 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3
* 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3
* 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5
* 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5
* 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5
* 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage
* 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage
* 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5
* 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]]
* 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5
* 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5
* 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3
* 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet
* 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance
* 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet
* 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance
* 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance
* 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie
* 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance
* 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3
* 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage
* 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm
* 07:29 moritzm: installing gnutls28 security updates
* 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie
* 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie
* 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage
* 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage
* 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm
* 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage
* 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]]
* 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage
* 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie
* 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie
* 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie
* 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie
* 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage
* 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage
* 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage
* 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage
* 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage
* 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage
* 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie
* 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie
* 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie
* 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s)
* 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]]
* 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s)
* 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]]
* 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie
* 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie
* 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage
* 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage
* 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage
* 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage
* 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie
* 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie
== 2026-07-07 ==
* 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie
* 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage
* 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage
* 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas)
* 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie
* 22:07 hashar: Restarting Gerrit on gerrit2003
* 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie
* 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie
* 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie
* 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage
* 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage
* 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage
* 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage
* 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage
* 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet
* 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage
* 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet
* 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006""
* 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s)
* 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006""
* 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie
* 20:17 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie
* 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie
* 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]]
* 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space
* 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie
* 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie
* 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie
* 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie
* 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie
* 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie
* 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage
* 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage
* 19:17 jasmine@dns1004: END - running authdns-update
* 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage
* 19:15 jasmine@dns1004: START - running authdns-update
* 19:15 cdobbins@dns1004: END - running authdns-update
* 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage
* 19:13 cdobbins@dns1004: START - running authdns-update
* 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update
* 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage
* 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage
* 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie
* 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie
* 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie
* 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]]
* 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]]
* 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]]
* 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]]
* 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]]
* 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]]
* 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]]
* 17:43 swfrench@dns1004: END - running authdns-update
* 17:40 swfrench@dns1004: START - running authdns-update
* 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]]
* 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet
* 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet
* 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet
* 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet
* 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]]
* 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet
* 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie
* 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]]
* 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet
* 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm
* 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]]
* 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]]
* 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage
* 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage
* 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794
* 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794
* 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie
* 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie
* 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage
* 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie
* 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage
* 15:54 mutante: jenkins down in planned maintenance window
* 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet
* 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet
* 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet
* 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie
* 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage
* 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm
* 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie
* 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage
* 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org
* 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org
* 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121
* 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121
* 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org
* 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121
* 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply
* 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003"
* 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003"
* 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie
* 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox
* 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org
* 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121
* 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org
* 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org
* 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie
* 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie
* 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org
* 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org
* 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s)
* 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]]
* 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s)
* 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]]
* 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s)
* 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f]
* 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s)
* 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage
* 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f]
* 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s)
* 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f]
* 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance
* 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage
* 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table
* 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage
* 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage
* 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie
* 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad
* 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037
* 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037
* 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]]
* 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s)
* 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037
* 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad
* 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]]
* 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie
* 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99)
* 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]]
* 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003"
* 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003"
* 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage
* 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox
* 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw
* 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037
* 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie
* 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet
* 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet
* 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet
* 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie
* 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage
* 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet
* 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet
* 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet
* 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw
* 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie
* 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases
* 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040
* 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s)
* 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage
* 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage
* 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts
* 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply
* 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage
* 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage
* 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply
* 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:38 moritzm: installing e2fsprogs updates from Trixie point release
* 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]]
* 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie
* 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade
* 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie
* 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
* 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.*
* 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie
* 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage
* 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie
* 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage
* 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]]
* 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie
* 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G
* 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage
* 13:13 jmm@dns1004: END - running authdns-update
* 13:12 jmm@dns1004: START - running authdns-update
* 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage
* 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie
* 13:07 jmm@dns1004: END - running authdns-update
* 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm
* 13:05 jmm@dns1004: START - running authdns-update
* 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage
* 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]]
* 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage
* 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin
* 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin
* 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage
* 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie
* 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage
* 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage
* 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003"
* 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003"
* 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie
* 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage
* 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036
* 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036
* 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie
* 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie
* 12:30 jmm@dns1004: END - running authdns-update
* 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036
* 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003"
* 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003"
* 12:28 jmm@dns1004: START - running authdns-update
* 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox
* 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036
* 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie
* 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet
* 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie
* 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet
* 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet
* 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]]
* 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter
* 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage
* 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage
* 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad
* 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet
* 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet
* 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet
* 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet
* 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad
* 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie
* 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence
* 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]]
* 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin
* 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin
* 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet
* 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie
* 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet
* 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie
* 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage
* 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage
* 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie
* 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage
* 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage
* 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage
* 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie
* 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie
* 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s)
* 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage
* 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]]
* 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s)
* 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie
* 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]]
* 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004
* 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie
* 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004
* 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003"
* 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003"
* 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot
* 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot
* 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet
* 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004
* 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie
* 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage
* 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage
* 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage
* 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage
* 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates
* 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie
* 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie
* 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates
* 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates
* 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates
* 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates
* 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates
* 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates
* 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:45 filippo@dns1006: END - running authdns-update
* 08:43 filippo@dns1006: START - running authdns-update
* 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]]
* 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates
* 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates
* 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet
* 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates
* 07:42 Msz2001: Deployed private patch for Suggested Ivestigations
* 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates
* 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates
* 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003"
* 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003
* 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates
* 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates
* 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003
* 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003"
* 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet
* 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates
* 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates
* 06:42 moritzm: install nginx security updates
* 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates
* 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates
* 06:19 moritzm: installing php8.2 security updates
* 06:15 moritzm: installing php8.4 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s)
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-06 ==
* 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s)
* 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie
* 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment
* 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug)
* 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]]
* 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage
* 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage
* 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie
* 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes
* 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes
* 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]]
* 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie
* 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie
* 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie
* 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage
* 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage
* 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s)
* 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage
* 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment
* 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage
* 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage
* 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]]
* 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie
* 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage
* 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie
* 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie
* 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie
* 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage
* 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage
* 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie
* 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie
* 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie
* 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage
* 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage
* 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage
* 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie
* 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage
* 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie
* 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie
* 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie
* 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie
* 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage
* 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage
* 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage
* 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage
* 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie
* 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie
* 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie
* 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie
* 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie
* 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie
* 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage
* 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage
* 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage
* 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage
* 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm
* 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage
* 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage
* 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie
* 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie
* 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie
* 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts
* 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie
* 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie
* 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s)
* 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie
* 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage
* 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage
* 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage
* 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage
* 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage
* 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie
* 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie
* 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm
* 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]]
* 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]]
* 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]]
* 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet
* 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet
* 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet
* 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97)
* 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]]
* 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw
* 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet
* 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet
* 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet
* 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet
* 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet
* 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet
* 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet
* 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet
* 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw
* 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s)
* 12:24 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet
* 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]]
* 11:57 moritzm: installing curl security updates
* 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:31 moritzm: installing nano security updates
* 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]]
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:50 jmm@dns1004: END - running authdns-update
* 10:47 jmm@dns1004: START - running authdns-update
* 10:47 jmm@dns1004: START - running authdns-update
* 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage
* 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage
* 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]]
* 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:24 elukey: spicerack 13.0.0 deployed on cumin2002
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie
* 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]]
* 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4
* 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6
* 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet
* 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]]
* 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet
* 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s)
* 07:57 fabfur: repooled cp4038
* 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.*
* 07:53 moritzm: installing pyjwt security updates
* 07:47 moritzm: installing openjpeg2 security updates
* 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment
* 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:38 moritzm: installing python-urllib3 security updates
* 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure
* 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.*
* 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.*
* 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]]
* 06:13 moritzm: installing Linux 6.12.95 on trixie hosts
* 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6
* 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4
* 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-05 ==
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-04 ==
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-03 ==
* 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade
* 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo
* 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo
* 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]]
* 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]]
* 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo
* 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo
* 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 14:40 cmooney@dns3003: END - running authdns-update
* 14:26 cmooney@dns3003: START - running authdns-update
* 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003"
* 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003"
* 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:38 sukhe@dns1004: END - running authdns-update
* 13:35 sukhe@dns1004: START - running authdns-update
* 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet
* 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]]
* 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet
* 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org
* 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org
* 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet
* 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003"
* 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003"
* 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox
* 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet
* 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet
* 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet
* 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]]
* 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org
* 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org
* 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-02 ==
* 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie
* 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage
* 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage
* 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie
* 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]]
* 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage
* 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s)
* 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards
* 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]]
* 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003
* 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003
* 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s)
* 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment
* 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s)
* 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie
* 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]]
* 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s)
* 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage
* 20:13 sbassett@deploy1003: sbassett: Continuing with deployment
* 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]]
* 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage
* 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards
* 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie
* 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]]
* 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards
* 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm
* 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet
* 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet
* 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet
* 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005"
* 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s)
* 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage
* 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]]
* 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage
* 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts
* 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003"
* 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003"
* 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003
* 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003
* 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm
* 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003"
* 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003"
* 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards
* 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm
* 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
* 16:57 tappof: bump space for prometheus k8s-aux in eqiad
* 16:55 cmooney@dns3003: END - running authdns-update
* 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003"
* 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003"
* 16:53 cmooney@dns3003: START - running authdns-update
* 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure)
* 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
* 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 16:39 rzl@dns1004: END - running authdns-update
* 16:37 rzl@dns1004: START - running authdns-update
* 16:36 rzl@dns1004: START - running authdns-update
* 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s)
* 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage
* 16:30 rzl@deploy1003: rzl: Continuing with deployment
* 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage
* 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]]
* 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync
* 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync
* 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm
* 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm
* 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates
* 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates
* 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage
* 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage
* 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates
* 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates
* 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm
* 15:24 moritzm: installing busybox updates from bookworm point release
* 15:20 moritzm: installing busybox updates from trixie point release
* 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates
* 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates
* 15:13 moritzm: installing giflib security updates
* 15:08 moritzm: installing Tomcat security updates
* 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003"
* 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates
* 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates
* 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003
* 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003"
* 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json
* 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json
* 14:32 moritzm: installing libdbi-perl security updates
* 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json
* 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json
* 14:12 moritzm: installing rsync security updates
* 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json
* 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance
* 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover
* 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:06 Tran: Deployed patch for [[phab:T427287|T427287]]
* 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover
* 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover
* 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover
* 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json
* 13:54 moritzm: installing sed security updates
* 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json
* 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]]
* 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json
* 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]]
* 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough
* 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org
* 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough
* 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough
* 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough
* 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s)
* 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment
* 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]]
* 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet
* 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 12:10 jmm@dns1004: END - running authdns-update
* 12:07 jmm@dns1004: START - running authdns-update
* 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox
* 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet
* 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling
* 10:49 jmm@dns1004: END - running authdns-update
* 10:47 jmm@dns1004: START - running authdns-update
* 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json
* 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json
* 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet
* 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]]
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003"
* 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003"
* 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet
* 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet
* 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet
* 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling
* 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json
* 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox
* 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet
* 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission
* 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json
* 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json
* 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance
* 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover
* 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json
* 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json
* 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]]
* 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json
* 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json
* 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]]
* 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json
* 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s)
* 09:13 moritzm: installing libgcrypt20 security updates
* 09:12 kharlan@deploy1003: kharlan: Continuing with deployment
* 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json
* 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]]
* 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s)
* 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json
* 08:57 kharlan@deploy1003: kharlan: Continuing with deployment
* 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]]
* 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json
* 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance
* 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s)
* 08:21 cscott@deploy1003: cscott: Continuing with deployment
* 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]]
* 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed
* 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s)
* 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org
* 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:57 cscott@deploy1003: cscott: Continuing with deployment
* 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org
* 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org
* 07:44 moritzm: installing node-lodash security updates
* 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]]
* 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org
* 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s)
* 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed
* 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]]
* 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s)
* 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie
* 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment
* 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]]
* 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage
* 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage
* 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie
* 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie
* 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet
* 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet
* 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]]
* 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage
* 06:14 cwilliams@dns1006: END - running authdns-update
* 06:12 cwilliams@dns1006: START - running authdns-update
* 06:11 cwilliams@dns1006: END - running authdns-update
* 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json
* 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage
* 06:09 cwilliams@dns1006: START - running authdns-update
* 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json
* 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json
* 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]]
* 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
* 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json
* 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]]
* 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie
* 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]]
* 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]]
* 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]]
* 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s)
* 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment
* 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed
* 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie
* 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration
* 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage
* 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage
* 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie
* 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111`
== 2026-07-01 ==
* 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election
* 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm
* 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie
* 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage
* 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage
* 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003
* 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003"
* 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003"
* 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage
* 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003
* 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm
* 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie
* 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie
* 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie
* 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage
* 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage
* 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie
* 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage
* 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage
* 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie
* 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie
* 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002"
* 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002"
* 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie
* 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage
* 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage
* 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage
* 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage
* 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie
* 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie
* 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s)
* 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment
* 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]]
* 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts
* 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts
* 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet
* 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet
* 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt
* 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt
* 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet
* 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet
* 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie
* 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]]
* 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts
* 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts
* 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie
* 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts
* 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s)
* 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage
* 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage
* 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage
* 16:18 jasmine@dns1004: END - running authdns-update
* 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage
* 16:15 jasmine@dns1004: START - running authdns-update
* 16:14 jasmine@dns1004: END - running authdns-update
* 16:12 jasmine@dns1004: START - running authdns-update
* 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance
* 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde
* 16:00 papaul: ongoing maintenance on lsw1-b2-codfw
* 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie
* 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt
* 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt
* 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie
* 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover
* 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie
* 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]]
* 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]]
* 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage
* 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:12 _joe_: restarted manually alertmanager-irc-relay
* 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage
* 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde
* 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover
* 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:02 papaul: ongoing maintenance on lsw1-a8-codfw
* 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]]
* 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s)
* 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]]
* 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover
* 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover
* 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]]
* 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json
* 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json
* 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment
* 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]]
* 14:03 jmm@dns1004: END - running authdns-update
* 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]]
* 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad
* 14:01 jmm@dns1004: START - running authdns-update
* 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json
* 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]]
* 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]]
* 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]]
* 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org
* 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie
* 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org
* 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie
* 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org
* 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org
* 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad
* 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad
* 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s)
* 13:30 caro@deploy1003: caro: Continuing with deployment
* 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:27 topranks: route-engine failover cr1-eqiad
* 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]]
* 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]]
* 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s)
* 13:13 moritzm: installing qemu security updates
* 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment
* 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance
* 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover
* 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]]
* 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover
* 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json
* 13:01 moritzm: installing python3.13 security updates
* 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json
* 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]]
* 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install
* 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json
* 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]]
* 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet
* 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet
* 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie
* 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie
* 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage
* 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed
* 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage
* 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage
* 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage
* 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]]
* 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie
* 11:50 cmooney@dns2005: END - running authdns-update
* 11:49 cmooney@dns2005: START - running authdns-update
* 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie
* 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed
* 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie
* 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie
* 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie
* 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie
* 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage
* 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage
* 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage
* 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage
* 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage
* 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie
* 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage
* 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage
* 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet
* 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet
* 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]]
* 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage
* 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json
* 10:26 moritzm: installing nginx security updates
* 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie
* 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json
* 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]]
* 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie
* 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie
* 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json
* 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]]
* 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003"
* 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003
* 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003
* 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003"
* 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s)
* 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s)
* 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment
* 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]]
* 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s)
* 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050]
* 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s)
* 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s)
* 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050]
* 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s)
* 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment
* 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050]
* 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s)
* 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]]
* 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050]
* 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s)
* 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment
* 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]]
* 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]]
* 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts
* 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie
* 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues
* 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet
* 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie
* 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage
* 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2
* 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7
* 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage
* 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage
* 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie
* 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage
* 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie
* 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie
* 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage
* 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage
* 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie
* 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash
* 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie
* 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage
* 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie
* 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie
* 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage
* 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage
* 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie
* 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage
* 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage
* 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage
* 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s)
* 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json
* 01:05 ladsgroup@dns1004: END - running authdns-update
* 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json
* 01:03 ladsgroup@dns1004: START - running authdns-update
* 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json
* 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]]
* 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json
* 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]]
* 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json
* 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie
* 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie
* 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally)
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
6xh269xx2v9n7z622sm6x2jpo1hxi2z
2445247
2445246
2026-08-09T02:07:56Z
Stashbot
7414
mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s)
2445247
wikitext
text/x-wiki
== 2026-08-09 ==
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-08 ==
* 05:31 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 02m 36s)
* 05:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9] (wcqs): [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry)
* 04:56 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 04:55 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 04:47 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry) (duration: 19m 22s)
* 04:28 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: deploy 0.3.165 (response_size telemetry)
* 04:19 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 06s)
* 04:18 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry)
* 04:17 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry) (duration: 00m 28s)
* 04:16 ryankemper@deploy1003: Started deploy [wdqs/wdqs@cf64fe9]: [[phab:T432775|T432775]]: 0.3.165 canary (response_size telemetry)
* 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 03:52 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-07 ==
* 23:30 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:29 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:45 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:43 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:41 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 20:54 jhancock@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-ctrl2006.codfw.wmnet with OS trixie
* 20:32 andrewbogott: restarting puppetserver service on puppetserver* for [[phab:T434339|T434339]]
* 19:52 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:45 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 19:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 19:32 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:25 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 19:22 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 19:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 18:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:41 jhancock@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 18:35 jhancock@cumin1003: START - Cookbook sre.hosts.provision for host wikikube-ctrl2006.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 18:23 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:22 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 18:22 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 18:16 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 18:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:09 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:08 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:07 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:06 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:04 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:01 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 18:00 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 17:59 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 17:25 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 17:14 ihurbain@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 17:13 ihurbain@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 17:13 ihurbain@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 17:13 ihurbain@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:54 maryum: Deployed security fix for [[phab:T434278|T434278]]
* 16:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:50 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 16:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 16:27 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:revalidateLinkRecommendations.php --wiki=enwiki --verbose --scoreLessThan=0.7 --exceptDatasetChecksums=[[phab:T434319|T434319]]-enwiki-models.txt # [[phab:T434319|T434319]]
* 16:06 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cp5022.eqsin.wmnet with OS trixie
* 15:13 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 14:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 14:17 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.clone (exit_code=0) of db1156.eqiad.wmnet onto db1271.eqiad.wmnet
* 13:50 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1271: Pool db1271.eqiad.wmnet in after cloning
* 13:02 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1271: Pool db1271.eqiad.wmnet in after cloning
* 12:19 jayme: updated calico to v3.30.7 on staging-codfw - [[phab:T427400|T427400]]
* 12:09 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:06 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:06 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'sync'.
* 12:05 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'sync'.
* 12:02 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Pool db1156.eqiad.wmnet in after cloning
* 11:38 bjensen: sudo -i reprepro -C main include trixie-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/trixie/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1+deb13u1_amd64.changes #[[phab:T434052|T434052]]
* 11:35 bjensen: sudo -i reprepro -C main include bookworm-wikimedia $<nowiki>{</nowiki>HOME<nowiki>}</nowiki>/httpbb/bookworm/httpbb_$<nowiki>{</nowiki>VERSION?<nowiki>}</nowiki>-1_amd64.changes #[[phab:T434052|T434052]]
* 11:30 marostegui@cumin1003: dbctl commit (dc=all): 'Adding db1271 to dbctl', diff saved to https://phabricator.wikimedia.org/P95945 and previous config saved to /var/cache/conftool/dbconfig/20260807-113006-marostegui.json
* 11:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Pool db1156.eqiad.wmnet in after cloning
* 10:23 jayme@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 10:22 jayme@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 10:22 jayme@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 10:21 jayme@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 10:20 jayme@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 10:20 jayme@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 10:19 jayme@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 10:18 jayme@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 10:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on 21 hosts with reason: cloning
* 10:01 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1156: Depool db1156.eqiad.wmnet to then clone it to db1271.eqiad.wmnet - marostegui@cumin1003
* 09:59 marostegui@cumin1003: START - Cookbook sre.mysql.clone of db1156.eqiad.wmnet onto db1271.eqiad.wmnet
* 09:15 jynus: started stress testing db1245 dbs [[phab:T431115|T431115]]
* 08:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:16 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:14 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:13 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:10 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:05 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:00 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 10 days, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows
* 08:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 07:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 07:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:49 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:47 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:45 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:37 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 06:35 jayme: updated istio to 1.29.4 on wikikube eqiad - [[phab:T427401|T427401]]
* 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1178.eqiad.wmnet
* 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 06:08 marostegui@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 06:06 marostegui@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1178.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - marostegui@cumin1003"
* 05:55 marostegui@cumin1003: START - Cookbook sre.dns.netbox
* 05:49 marostegui@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1178.eqiad.wmnet
* 05:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 05:46 marostegui@cumin1003: Removing db1178 from zarcillo [[phab:T433471|T433471]]
* 05:45 marostegui@cumin1003: START - Cookbook sre.mysql.decommission
* 02:42 denisse: Extended volume on prometheus2008 for the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space
* 02:37 denisse: Extended volume on prometheus2007 tor the disk space alert as per https://wikitech.wikimedia.org/wiki/Prometheus#Prometheus_host_running_out_of_space
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 56s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-06 ==
* 21:39 maryum: Deploy security patch for [[phab:T433070|T433070]]
* 21:29 maryum: Deploy security patch for [[phab:T434189|T434189]]
* 20:48 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] (duration: 08m 12s)
* 20:44 aude@deploy1003: lmora, aude, anzx: Continuing with deployment
* 20:41 aude@deploy1003: lmora, aude, anzx: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be
* 20:41 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:41 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:40 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1322038{{!}}outreachwiki: disable bureaucrats ability to locally remove users from importer usergroup (T431959)]], [[gerrit:1321639{{!}}Turn on feature flag for custom lists for betawiki (T434027)]], [[gerrit:1322054{{!}}tcywiki: update logos for 10years anniversary (T434176)]]
* 20:37 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:37 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:32 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:32 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:31 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] (duration: 06m 41s)
* 20:27 cjming@deploy1003: cjming, ebernhardson, chlod: Continuing with deployment
* 20:26 cjming@deploy1003: cjming, ebernhardson, chlod: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:24 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1319939{{!}}Revert "frwiki: change to Wikipedia 25 logo" (T430409)]], [[gerrit:1320230{{!}}cirrus: Enable building redirect documents (T204089)]]
* 20:18 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] (duration: 09m 22s)
* 20:14 cjming@deploy1003: cjming, tsev: Continuing with deployment
* 20:11 cjming@deploy1003: cjming, tsev: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1315128{{!}}Add new Apple app site association file for Test Wiki (T432412)]]
* 19:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:31 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 19:00 cdobbins@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cp5022.eqsin.wmnet with OS trixie
* 18:25 ladsgroup@deploy1003: Finished scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]]) (duration: 06m 08s)
* 18:19 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]])
* 18:18 ladsgroup@deploy1003: Stopping before sync operations
* 18:17 ladsgroup@deploy1003: Started scap sync-world: Deploying gerrit:1321603 ([[phab:T107188|T107188]])
* 17:55 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host cp5022.eqsin.wmnet with OS trixie
* 16:50 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 16:35 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 16:19 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage
* 16:16 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1002.eqiad.wmnet with reason: host reimage
* 16:05 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage
* 16:00 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1001.eqiad.wmnet with reason: host reimage
* 15:57 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 15:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 15:29 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:58 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 14:57 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs ([[phab:T428495|T428495]])
* 14:55 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs ([[phab:T428495|T428495]])
* 14:55 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-ui1001.eqiad.wmnet with OS bookworm
* 14:54 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-presto1001.eqiad.wmnet with OS bookworm
* 14:52 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru ([[phab:T428495|T428495]])
* 14:49 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru ([[phab:T428495|T428495]])
* 14:48 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 14:46 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams ([[phab:T428495|T428495]])
* 14:44 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams ([[phab:T428495|T428495]])
* 14:43 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:42 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:42 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:42 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 14:40 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 14:40 brouberol@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:38 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage
* 14:34 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage
* 14:28 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-ui1001.eqiad.wmnet with reason: host reimage
* 14:27 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-presto1001.eqiad.wmnet with reason: host reimage
* 14:23 sukhe: sudo cumin -b2 'A:cp-text' "run-puppet-agent --enable 'merging CR 1290731'": [[phab:T425441|T425441]]
* 14:18 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]]
* 14:16 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-presto1001.eqiad.wmnet with OS bookworm
* 14:14 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-ui1001.eqiad.wmnet with OS bookworm
* 14:12 sukhe: sudo cumin 'A:cp-text' "disable-puppet 'merging CR 1290731'": [[phab:T425441|T425441]]
* 14:11 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]]
* 14:04 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]]
* 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1001.eqiad.wmnet with OS bookworm
* 14:04 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1002.eqiad.wmnet with OS bookworm
* 14:02 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-client1002.eqiad.wmnet with OS bookworm
* 13:58 swfrench-wmf: authdns update to direct eqiad-associated etcd clients back to eqiad - [[phab:T428495|T428495]]
* 13:58 swfrench@dns1004: END - running authdns-update
* 13:56 swfrench@dns1004: START - running authdns-update
* 13:49 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:44 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:31 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 13:29 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 13:26 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 13:23 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 13:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 13:19 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 13:18 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 13:18 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 13:17 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 13:16 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 13:13 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 13:11 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 13:09 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 13:06 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 13:06 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm
* 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' .
* 13:05 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:05 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revision-models' for release 'main' .
* 13:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host an-test-client1002.eqiad.wmnet with OS bookworm
* 13:04 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 13:03 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 13:02 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 13:00 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'readability' for release 'main' .
* 12:59 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'readability' for release 'main' .
* 12:58 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' .
* 12:57 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'logo-detection' for release 'main' .
* 12:57 aikochou@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 12:55 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:54 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:53 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-client1002.eqiad.wmnet with reason: host reimage
* 12:53 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:50 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
* 12:48 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
* 12:46 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
* 12:45 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-models' for release 'main' .
* 12:41 bwojtowicz@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' .
* 12:39 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'article-descriptions' for release 'main' .
* 12:38 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-client1002.eqiad.wmnet with OS bookworm
* 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 12:12 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 12:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 11:58 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2187: Security update
* 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:31 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:26 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:24 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-pageview: apply
* 11:16 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:15 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 11:10 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2187: Security update
* 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2187.codfw.wmnet with reason: Maintenance
* 10:56 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 10:56 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 10:56 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 10:56 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 10:54 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 10:53 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 10:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2187: Security update
* 10:07 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2187: Security update
* 09:39 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms2', diff saved to https://phabricator.wikimedia.org/P95929 and previous config saved to /var/cache/conftool/dbconfig/20260806-093908-marostegui.json
* 09:36 marostegui@cumin1003: dbctl commit (dc=all): 'Remove db1178 from dbctl [[phab:T433471|T433471]]', diff saved to https://phabricator.wikimedia.org/P95928 and previous config saved to /var/cache/conftool/dbconfig/20260806-093632-marostegui.json
* 09:33 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 09:31 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 09:30 topranks: bounce cr3-eqsin<->cr2-eqiad bgp session to disable no-prepend command
* 09:20 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2253.codfw.wmnet,db1151.eqiad.wmnet with reason: cloning
* 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Cloning
* 09:20 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:19 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Cloning
* 09:10 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 09:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup2003.codfw.wmnet
* 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:07 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 09:06 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup2003.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 09:03 klausman@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:02 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 09:02 klausman@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:57 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup2003.codfw.wmnet
* 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup1003.eqiad.wmnet
* 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:55 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:54 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms3', diff saved to https://phabricator.wikimedia.org/P95925 and previous config saved to /var/cache/conftool/dbconfig/20260806-085422-marostegui.json
* 08:53 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup1003.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:46 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 08:39 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup1003.eqiad.wmnet
* 08:29 XioNoX: push pfw policy - [[phab:T434115|T434115]]
* 08:14 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 08:00 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:58 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revision-models' for release 'main' .
* 07:56 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 07:54 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'readability' for release 'main' .
* 07:53 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'logo-detection' for release 'main' .
* 07:51 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'llm' for release 'main' .
* 07:48 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-descriptions' for release 'main' .
* 07:37 jayme: updated istio to 1.29.4 on wikikube codfw - [[phab:T427401|T427401]]
* 07:08 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2252.codfw.wmnet,db1153.eqiad.wmnet with reason: cloning
* 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Cloning
* 07:07 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:06 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Cloning
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 40s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-05 ==
* 23:24 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1009.eqiad.wmnet with OS bookworm
* 23:03 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage
* 22:59 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1009.eqiad.wmnet with reason: host reimage
* 22:43 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1009.eqiad.wmnet with OS bookworm
* 22:38 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1009.eqiad.wmnet
* 22:34 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1009.eqiad.wmnet
* 22:25 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1008.eqiad.wmnet with OS bookworm
* 22:04 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage
* 22:00 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1008.eqiad.wmnet with reason: host reimage
* 21:48 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:47 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:46 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:44 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1008.eqiad.wmnet with OS bookworm
* 21:43 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 21:41 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1008.eqiad.wmnet
* 21:36 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1008.eqiad.wmnet
* 21:14 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad
* 21:12 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad
* 21:10 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad
* 21:08 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:07 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:07 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad
* 21:04 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 21:03 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006
* 21:02 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006
* 21:00 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:56 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:56 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:55 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 20:55 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 20:55 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf1007.eqiad.wmnet with OS bookworm
* 20:51 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 20:43 ebernhardson: [[phab:T434008|T434008]]: changing cloudelastic:9643 from auto_expand_replicas to number_of_replicas
* 20:34 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage
* 20:27 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf1007.eqiad.wmnet with reason: host reimage
* 20:24 cjming: end of UTC late backport window
* 20:23 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] (duration: 06m 26s)
* 20:18 cjming@deploy1003: cjming: Continuing with deployment
* 20:18 cjming@deploy1003: cjming: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:16 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1321591{{!}}logging: Update wikilambda/ui_actions schema ref. from 1.0.0 to 1.1.0 (T433550)]]
* 20:12 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm
* 20:12 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] (duration: 08m 41s)
* 20:08 swfrench@cumin2002: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host conf1007.eqiad.wmnet with OS bookworm
* 20:08 jforrester@deploy1003: jforrester: Continuing with deployment
* 20:07 jforrester@deploy1003: jforrester: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:03 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1320196{{!}}ExtensionDistributor: Drop REL1_44, EOL (T428911)]], [[gerrit:1311106{{!}}wikifunctions: Configure wgWikiLambdaClientRepoSiteId so RC entries point correctly]], [[gerrit:1321585{{!}}abstractwiki: Add dag/ml/ig/ha languages to generation script]]
* 19:51 inflatador: [bking@puppetserver1001] ~$ sudo puppetserver ca sign --certname an-worker1189.eqiad.wmnet [[phab:T434142|T434142]]
* 19:47 bking@cumin2003: DONE (FAIL) - Cookbook sre.puppet.renew-cert (exit_code=99) for an-worker1189.eqiad.wmnet: Renew puppet certificate - bking@cumin2003
* 19:46 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1007.eqiad.wmnet with OS trixie
* 19:30 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 19:29 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 19:20 swfrench-wmf: silenced EtcdRelicationDown 0cb709a9-f244-4f1e-971f-{{Gerrit|440ec65e7fd7}} - [[phab:T428495|T428495]]
* 19:13 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf1007.eqiad.wmnet with OS bookworm
* 19:12 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage
* 19:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf1007.eqiad.wmnet
* 19:07 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1007.eqiad.wmnet with reason: host reimage
* 19:03 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf1007.eqiad.wmnet
* 18:52 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1007.eqiad.wmnet with OS trixie
* 18:52 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:35 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 18:34 fceratto@deploy1003: helmfile [aux-k8s-eqiad] 'sync' command on namespace 'zarcillo' for release 'main' .
* 18:30 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1007.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:28 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003"
* 18:27 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1007] - vriley@cumin1003"
* 18:23 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:22 vriley@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 18:22 robh@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:19 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 18:13 robh@cumin2002: START - Cookbook sre.hosts.provision for host cp5022.mgmt.eqsin.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:31 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-eqiad
* 17:12 mutante: LDAP - added vwalters to group ciadmin - [[phab:T433615|T433615]]
* 16:58 aokoth@deploy1003: Finished deploy [phabricator/deployment@e2ebca5]: Deploy Phab (duration: 00m 34s)
* 16:57 aokoth@deploy1003: Started deploy [phabricator/deployment@e2ebca5]: Deploy Phab
* 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad
* 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad
* 16:55 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad
* 16:53 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2187.codfw.wmnet
* 16:41 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2187.codfw.wmnet
* 16:41 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2187.codfw.wmnet
* 16:41 cgoubert@cumin2003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2187.codfw.wmnet
* 16:40 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-eqiad
* 16:40 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:34 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:25 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]])
* 16:23 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-magru and A:liberica ([[phab:T428495|T428495]])
* 16:20 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]])
* 16:19 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-drmrs and A:liberica ([[phab:T428495|T428495]])
* 16:18 cdobbins@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]])
* 16:16 cdobbins@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-esams and A:liberica ([[phab:T428495|T428495]])
* 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1159.eqiad.wmnet
* 16:06 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1159.eqiad.wmnet
* 16:06 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1159.eqiad.wmnet
* 16:05 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] (duration: 09m 11s)
* 15:58 reedy@deploy1003: reedy: Continuing with deployment
* 15:58 reedy@deploy1003: reedy: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:56 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321583{{!}}maintenance: Add require_once statements for AllUsers (T420792)]]
* 15:54 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1159.eqiad.wmnet with OS trixie
* 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage
* 15:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:27 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1159.eqiad.wmnet with reason: host reimage
* 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1159
* 15:10 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1159
* 15:00 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo for hosts in the wikimedia.org domain - [[phab:T428495|T428495]]
* 14:55 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie
* 14:54 swfrench-wmf: restarted navtiming on webperf1003 - [[phab:T428495|T428495]]
* 14:52 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1159
* 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:52 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1159.eqiad.wmnet 129.48.64.10.in-addr.arpa 9.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:52 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003"
* 14:52 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1159 - jayme@cumin1003"
* 14:49 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:48 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 14:47 swfrench-wmf: begin rolling restart of confd in drmrs, eqiad, esams, magru - [[phab:T428495|T428495]]
* 14:47 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1159
* 14:46 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:46 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:46 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1159.eqiad.wmnet with OS trixie
* 14:45 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:44 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:44 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1159.eqiad.wmnet
* 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1159.eqiad.wmnet
* 14:43 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:43 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1159.eqiad.wmnet
* 14:43 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:43 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1157.eqiad.wmnet
* 14:43 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1157.eqiad.wmnet
* 14:43 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1157.eqiad.wmnet
* 14:42 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:42 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:42 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:41 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:41 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:41 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:41 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-worker1003.eqiad.wmnet with OS bookworm
* 14:39 swfrench-wmf: authdns update to direct eqiad-associated etcd clients to codfw - [[phab:T428495|T428495]]
* 14:39 swfrench@dns1004: END - running authdns-update
* 14:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 14:37 swfrench@dns1004: START - running authdns-update
* 14:37 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:37 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:35 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:34 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 14:28 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1157.eqiad.wmnet with OS trixie
* 14:27 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:27 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:27 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 14:26 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 14:26 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:26 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 14:26 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 14:26 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:26 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader1002.eqiad.wmnet with OS trixie
* 14:19 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:15 brouberol@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage
* 14:14 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046
* 14:13 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046
* 14:13 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie
* 14:11 brouberol@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-worker1003.eqiad.wmnet with reason: host reimage
* 14:10 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:09 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:09 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:09 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:08 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:08 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage
* 14:08 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:04 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_eqiad: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 14:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage
* 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=eqiad
* 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=eqiad
* 14:01 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=eqiad
* 14:00 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:59 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:58 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1157.eqiad.wmnet with reason: host reimage
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader1002.eqiad.wmnet with reason: host reimage
* 13:54 brouberol@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-worker1003.eqiad.wmnet with OS bookworm
* 13:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader1002.eqiad.wmnet with OS trixie
* 13:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1157
* 13:42 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1157
* 13:40 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1157
* 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1157.eqiad.wmnet 183.32.64.10.in-addr.arpa 3.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:40 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003"
* 13:39 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1157 - jayme@cumin1003"
* 13:39 reedy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] (duration: 07m 00s)
* 13:35 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 13:35 reedy@deploy1003: reedy: Continuing with deployment
* 13:34 reedy@deploy1003: reedy: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:32 reedy@deploy1003: Started scap sync-world: Backport for [[gerrit:1321547{{!}}wmf-config: OATHAuth config changes (T428103)]]
* 13:23 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1157
* 13:22 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1157.eqiad.wmnet with OS trixie
* 13:22 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1157.eqiad.wmnet
* 13:22 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1157.eqiad.wmnet
* 13:21 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1157.eqiad.wmnet
* 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1156.eqiad.wmnet
* 13:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1156.eqiad.wmnet
* 13:15 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1156.eqiad.wmnet
* 13:01 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1156.eqiad.wmnet with OS trixie
* 12:42 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage
* 12:38 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1156.eqiad.wmnet with reason: host reimage
* 12:32 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 12:31 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 12:30 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 12:28 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 12:26 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 12:24 topranks: update bgp confed settings in eqsin
* 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1156
* 12:22 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1156
* 12:22 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 12:19 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1156
* 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:19 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1156.eqiad.wmnet 110.32.64.10.in-addr.arpa 0.1.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:19 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003"
* 12:19 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1156 - jayme@cumin1003"
* 12:17 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:14 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS trixie
* 12:09 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:06 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:04 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 12:04 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'articletopic-outlink' for release 'main' .
* 12:02 bwojtowicz@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'article-models' for release 'main' .
* 12:01 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1156
* 12:01 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1156.eqiad.wmnet with OS trixie
* 11:59 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1156.eqiad.wmnet
* 11:59 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1156.eqiad.wmnet
* 11:59 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1156.eqiad.wmnet
* 11:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 11:53 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 11:53 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:52 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:50 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:50 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:48 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:47 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:47 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:45 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:45 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:44 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:44 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:43 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:42 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:38 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:35 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046
* 11:35 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host mc2046
* 11:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS trixie
* 11:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:21 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:21 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:18 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:18 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:18 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:10 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:09 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:08 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:06 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:06 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:05 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 11:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 11:04 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:24 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:24 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:23 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:17 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1155.eqiad.wmnet
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1155.eqiad.wmnet
* 10:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1155.eqiad.wmnet
* 10:14 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:14 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:11 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:11 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:09 aikochou@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync
* 10:09 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:09 aikochou@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: sync
* 10:07 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:07 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 10:05 aikochou@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync
* 10:05 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:05 aikochou@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: sync
* 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:04 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 10:04 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 10:04 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1155.eqiad.wmnet with OS trixie
* 09:52 marostegui@cumin1003: dbctl commit (dc=all): 'Pool back ms1', diff saved to https://phabricator.wikimedia.org/P95918 and previous config saved to /var/cache/conftool/dbconfig/20260805-095212-marostegui.json
* 09:44 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: after cloning
* 09:44 marostegui@cumin1003: END (FAIL) - Cookbook sre.mysql.parsercache (exit_code=99)
* 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:44 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1152: after cloning
* 09:43 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage
* 09:40 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1155.eqiad.wmnet with reason: host reimage
* 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 09:32 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 09:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1155
* 09:27 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1155
* 09:25 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 09:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 09:24 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 09:23 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:23 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 09:23 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:22 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 09:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 09:22 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:20 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 09:20 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:17 XioNoX: push pfw policies - [[phab:T434038|T434038]]
* 09:14 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1155
* 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1155.eqiad.wmnet 109.32.64.10.in-addr.arpa 9.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:14 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003"
* 09:14 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1155 - jayme@cumin1003"
* 09:10 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 09:10 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:09 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: cloning
* 09:09 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 09:08 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 09:08 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Cloning
* 09:06 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Cloning
* 08:38 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1155
* 08:37 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1155.eqiad.wmnet with OS trixie
* 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1171.eqiad.wmnet
* 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:29 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95913 and previous config saved to /var/cache/conftool/dbconfig/20260805-082908-ladsgroup.json
* 08:27 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1171.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:22 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 08:18 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95912 and previous config saved to /var/cache/conftool/dbconfig/20260805-081823-ladsgroup.json
* 08:17 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1171.eqiad.wmnet
* 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts db1150.eqiad.wmnet
* 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:17 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:15 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: db1150.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 08:15 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 08:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1155.eqiad.wmnet
* 08:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1155.eqiad.wmnet
* 08:14 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1155.eqiad.wmnet
* 08:11 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 08:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249', diff saved to https://phabricator.wikimedia.org/P95911 and previous config saved to /var/cache/conftool/dbconfig/20260805-080737-ladsgroup.json
* 08:05 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts db1150.eqiad.wmnet
* 08:02 marostegui: Depool clouddb1020 (s5,s8) [[phab:T434048|T434048]]
* 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s8
* 08:02 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1020.eqiad.wmnet,service=s5
* 08:02 marostegui: Depool clouddb1018 (s2,s7) [[phab:T434048|T434048]]
* 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s7
* 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1018.eqiad.wmnet,service=s2
* 08:01 marostegui: Depool clouddb1017 (s1) [[phab:T434048|T434048]]
* 08:01 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1
* 07:59 marostegui: Depool clouddb1016 (s5,s8) [[phab:T434048|T434048]]
* 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8
* 07:59 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5
* 07:57 marostegui: Depool clouddb1015 (s4,s6) [[phab:T434048|T434048]]
* 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6
* 07:57 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4
* 07:56 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95910 and previous config saved to /var/cache/conftool/dbconfig/20260805-075650-ladsgroup.json
* 07:54 marostegui: Depool clouddb1014 (s2,s7) [[phab:T434048|T434048]]
* 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s7
* 07:54 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1014.eqiad.wmnet,service=s2
* 07:53 marostegui: Depool clouddb1013:s1 [[phab:T434048|T434048]]
* 07:53 marostegui: Depool clouddb1013:s1 [[phab:T409557|T409557]]
* 07:53 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1013.eqiad.wmnet,service=s1
* 07:25 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2249 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95909 and previous config saved to /var/cache/conftool/dbconfig/20260805-072529-ladsgroup.json
* 07:24 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2249.codfw.wmnet with reason: Maintenance
* 07:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95908 and previous config saved to /var/cache/conftool/dbconfig/20260805-072426-ladsgroup.json
* 07:21 slyngshede@dns1004: END - running authdns-update
* 07:19 slyngshede@dns1004: START - running authdns-update
* 07:13 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95906 and previous config saved to /var/cache/conftool/dbconfig/20260805-071340-ladsgroup.json
* 07:02 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231', diff saved to https://phabricator.wikimedia.org/P95905 and previous config saved to /var/cache/conftool/dbconfig/20260805-070253-ladsgroup.json
* 06:52 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95904 and previous config saved to /var/cache/conftool/dbconfig/20260805-065206-ladsgroup.json
* 06:45 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 06:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95903 and previous config saved to /var/cache/conftool/dbconfig/20260805-062240-ladsgroup.json
* 06:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2231.codfw.wmnet with reason: Maintenance
* 06:21 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95902 and previous config saved to /var/cache/conftool/dbconfig/20260805-062137-ladsgroup.json
* 06:10 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95901 and previous config saved to /var/cache/conftool/dbconfig/20260805-061051-ladsgroup.json
* 06:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215', diff saved to https://phabricator.wikimedia.org/P95900 and previous config saved to /var/cache/conftool/dbconfig/20260805-060004-ladsgroup.json
* 05:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95899 and previous config saved to /var/cache/conftool/dbconfig/20260805-054918-ladsgroup.json
* 05:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2215 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95898 and previous config saved to /var/cache/conftool/dbconfig/20260805-051939-ladsgroup.json
* 05:18 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2215.codfw.wmnet with reason: Maintenance
* 04:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2201.codfw.wmnet with reason: Maintenance
* 03:40 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2197.codfw.wmnet with reason: Maintenance
* 03:40 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95897 and previous config saved to /var/cache/conftool/dbconfig/20260805-034036-ladsgroup.json
* 03:29 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95896 and previous config saved to /var/cache/conftool/dbconfig/20260805-032948-ladsgroup.json
* 03:19 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196', diff saved to https://phabricator.wikimedia.org/P95895 and previous config saved to /var/cache/conftool/dbconfig/20260805-031902-ladsgroup.json
* 03:08 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95894 and previous config saved to /var/cache/conftool/dbconfig/20260805-030815-ladsgroup.json
* 02:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95893 and previous config saved to /var/cache/conftool/dbconfig/20260805-023413-ladsgroup.json
* 02:33 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2196.codfw.wmnet with reason: Maintenance
* 02:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95892 and previous config saved to /var/cache/conftool/dbconfig/20260805-023310-ladsgroup.json
* 02:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95891 and previous config saved to /var/cache/conftool/dbconfig/20260805-022223-ladsgroup.json
* 02:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186', diff saved to https://phabricator.wikimedia.org/P95890 and previous config saved to /var/cache/conftool/dbconfig/20260805-021137-ladsgroup.json
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 02:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95889 and previous config saved to /var/cache/conftool/dbconfig/20260805-020051-ladsgroup.json
* 01:30 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95888 and previous config saved to /var/cache/conftool/dbconfig/20260805-013029-ladsgroup.json
* 01:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db2186.codfw.wmnet with reason: Maintenance
* 00:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on dbstore1009.eqiad.wmnet with reason: Maintenance
* 00:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95887 and previous config saved to /var/cache/conftool/dbconfig/20260805-003408-ladsgroup.json
* 00:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95886 and previous config saved to /var/cache/conftool/dbconfig/20260805-002322-ladsgroup.json
* 00:12 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264', diff saved to https://phabricator.wikimedia.org/P95885 and previous config saved to /var/cache/conftool/dbconfig/20260805-001235-ladsgroup.json
* 00:01 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95884 and previous config saved to /var/cache/conftool/dbconfig/20260805-000148-ladsgroup.json
== 2026-08-04 ==
* 23:45 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95883 and previous config saved to /var/cache/conftool/dbconfig/20260804-234508-ladsgroup.json
* 23:44 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1264.eqiad.wmnet with reason: Maintenance
* 23:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95882 and previous config saved to /var/cache/conftool/dbconfig/20260804-234405-ladsgroup.json
* 23:33 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95881 and previous config saved to /var/cache/conftool/dbconfig/20260804-233317-ladsgroup.json
* 23:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237', diff saved to https://phabricator.wikimedia.org/P95880 and previous config saved to /var/cache/conftool/dbconfig/20260804-232230-ladsgroup.json
* 23:11 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95879 and previous config saved to /var/cache/conftool/dbconfig/20260804-231144-ladsgroup.json
* 22:23 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95878 and previous config saved to /var/cache/conftool/dbconfig/20260804-222345-ladsgroup.json
* 22:23 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1237.eqiad.wmnet with reason: Maintenance
* 21:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1225.eqiad.wmnet with reason: Maintenance
* 20:40 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] (duration: 24m 40s)
* 20:33 samtar@deploy1003: samtar, kineticpelagic: Continuing with deployment
* 20:28 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc2046.codfw.wmnet with OS bookworm
* 20:21 samtar@deploy1003: samtar, kineticpelagic: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:15 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1313827{{!}}rest: Add test server option to REST Sandbox for Wikipedia projects (T408816)]]
* 20:13 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 20:09 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on mc2046.codfw.wmnet with reason: host reimage
* 20:00 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1216.eqiad.wmnet with reason: Maintenance
* 20:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95877 and previous config saved to /var/cache/conftool/dbconfig/20260804-195957-ladsgroup.json
* 19:51 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host mc2046
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 19:50 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:50 jhancock@cumin2002: START - Cookbook sre.dns.wipe-cache mc2046.codfw.wmnet 120.16.192.10.in-addr.arpa 0.2.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:50 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002"
* 19:50 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host mc2046 - jhancock@cumin2002"
* 19:49 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95876 and previous config saved to /var/cache/conftool/dbconfig/20260804-194911-ladsgroup.json
* 19:46 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.move-vlan for host mc2046
* 19:45 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host mc2046.codfw.wmnet with OS bookworm
* 19:38 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203', diff saved to https://phabricator.wikimedia.org/P95875 and previous config saved to /var/cache/conftool/dbconfig/20260804-193825-ladsgroup.json
* 19:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95874 and previous config saved to /var/cache/conftool/dbconfig/20260804-192738-ladsgroup.json
* 19:02 mutante: gerrit ssh -p 29418 gerrit.wikimedia.org gerrit index changes {{Gerrit|1320979}}
* 18:20 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 18:18 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 18:14 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 18:14 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 18:13 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 18:10 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 18:08 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 18:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95872 and previous config saved to /var/cache/conftool/dbconfig/20260804-180721-ladsgroup.json
* 18:07 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 18:06 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1203.eqiad.wmnet with reason: Maintenance
* 18:06 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95871 and previous config saved to /var/cache/conftool/dbconfig/20260804-180618-ladsgroup.json
* 17:55 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95870 and previous config saved to /var/cache/conftool/dbconfig/20260804-175531-ladsgroup.json
* 17:55 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1154.eqiad.wmnet
* 17:55 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1154.eqiad.wmnet
* 17:55 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1154.eqiad.wmnet
* 17:50 swfrench@deploy1003: Finished scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] (duration: 04m 05s)
* 17:48 swfrench@deploy1003: swfrench: Continuing with deployment
* 17:46 swfrench@deploy1003: swfrench: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:45 swfrench@deploy1003: Started scap sync-world: Deploy helmfile-only msmtp timeout override cleanup - [[phab:T383047|T383047]]
* 17:44 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179', diff saved to https://phabricator.wikimedia.org/P95869 and previous config saved to /var/cache/conftool/dbconfig/20260804-174445-ladsgroup.json
* 17:34 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95868 and previous config saved to /var/cache/conftool/dbconfig/20260804-173359-ladsgroup.json
* 17:33 swfrench@deploy1003: Finished scap sync-world: Pick up new PHP production image (duration: 28m 32s)
* 17:28 aokoth@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on phab1005.eqiad.wmnet with reason: Puppet Failure
* 17:05 swfrench@deploy1003: Started scap sync-world: Pick up new PHP production image
* 17:00 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=97)
* 17:00 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 16:54 cgoubert@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on wikikube-worker2187.codfw.wmnet with reason: Hardware issue
* 16:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2187.codfw.wmnet
* 16:52 mutante: gerrit2003:/var/log/apache2# ln -s /srv/gerrit/site_path/review_site/logs/ gerrit
* 16:52 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2187.codfw.wmnet
* 16:48 mutante: gerrit2003 - moving old apache logfiles older than 60 days from /var/log/apache2 to /srv/gerrit/site_path/review_site/logs/old/
* 16:33 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:32 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 16:29 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 16:28 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 16:27 dzahn@cumin1003: END (PASS) - Cookbook sre.gerrit.restart-gerrit (exit_code=0) Restarting Gerrit on gerrit2003
* 16:27 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:27 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95867 and previous config saved to /var/cache/conftool/dbconfig/20260804-162736-ladsgroup.json
* 16:27 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 16:26 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance
* 16:26 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 16:25 mutante: restarting gerrit - dropped outdated RSA host key
* 16:25 dzahn@cumin1003: START - Cookbook sre.gerrit.restart-gerrit Restarting Gerrit on gerrit2003
* 16:24 ladsgroup@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95866 and previous config saved to /var/cache/conftool/dbconfig/20260804-162424-ladsgroup.json
* 16:24 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 16:23 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 16:22 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T433990|T433990]])', diff saved to https://phabricator.wikimedia.org/P95865 and previous config saved to /var/cache/conftool/dbconfig/20260804-162236-ladsgroup.json
* 16:21 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on db1179.eqiad.wmnet with reason: Maintenance
* 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf11u1 into component/php83 for bullseye-wikimedia
* 16:17 swfrench-wmf: reprepro include php8.3_8.3.33-1+wmf12u1 into component/php83 for bookworm-wikimedia
* 16:11 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 16:10 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 16:10 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 16:09 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 16:09 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 16:08 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 16:08 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:08 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:07 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 16:07 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 16:07 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 16:05 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
* 16:04 mutante: gerrit2002/gerrit1003/gerrit2003 - rm /etc/gerrit/ssh_host_rsa_key
* 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:59 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:59 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['cp5022.mgmt.eqsin.wmnet']
* 15:56 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:55 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:55 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:49 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 15:49 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:44 Raine: add php8.5 packages to component/php85 - [[phab:T432983|T432983]]
* 15:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:33 aaron@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 15:33 aaron@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 15:29 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:19 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:19 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:16 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:16 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1154.eqiad.wmnet with OS trixie
* 15:16 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 15:15 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 15:06 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]] (duration: 00m 43s)
* 15:05 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab1004 for [[phab:T433981|T433981]]
* 15:05 aaron@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 15:04 aaron@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 15:02 brennen@deploy1003: Finished deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]] (duration: 00m 51s)
* 15:01 brennen@deploy1003: Started deploy [phabricator/deployment@56f4ffd]: deploy phab2003 for [[phab:T433981|T433981]]
* 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment
* 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment
* 14:58 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment
* 14:55 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage
* 14:51 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1154.eqiad.wmnet with reason: host reimage
* 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:48 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:38 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
* 14:38 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
* 14:38 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
* 14:37 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
* 14:37 ottomata: roll restart eventgate-main to pick up stream config change - [[phab:T433507|T433507]]
* 14:37 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventgate-main: sync
* 14:36 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventgate-main: sync
* 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1154
* 14:36 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1154
* 14:34 otto@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] (duration: 08m 39s)
* 14:34 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1154
* 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:34 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1154.eqiad.wmnet 108.32.64.10.in-addr.arpa 8.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:34 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003"
* 14:34 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1154 - jayme@cumin1003"
* 14:30 otto@deploy1003: otto: Continuing with deployment
* 14:30 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 14:28 otto@deploy1003: otto: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:26 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1154
* 14:26 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1154.eqiad.wmnet with OS trixie
* 14:26 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet
* 14:26 otto@deploy1003: Started scap sync-world: Backport for [[gerrit:1320941{{!}}EventStreamConfig - page_html_feature_counts_change.v1 canary to eventgate-main (T433507)]]
* 14:26 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet
* 14:26 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet
* 14:17 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 14:16 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 14:15 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 14:14 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 14:13 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 14:13 swfrench@dns1004: END - running authdns-update
* 14:13 Msz2001: Finished deployments for UTC afternoon backport window
* 14:13 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 14:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] (duration: 07m 58s)
* 14:11 swfrench@dns1004: START - running authdns-update
* 14:08 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Continuing with deployment
* 14:07 mszwarc@deploy1003: javiermonton, mszwarc, mpostoronca: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]] synced to the testser
* 14:05 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320757{{!}}stream: pageview.trending.relative.v1 (T432204)]], [[gerrit:1319086{{!}}WikimediaAntiAbuse: Document required load order after Echo (T432452)]], [[gerrit:1320781{{!}}UIC: Add user name to server-side instrumentation events (T433816)]], [[gerrit:1320782{{!}}UIC: Add user name to server-side instrumentation events (T433816)]]
* 14:03 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (3 nodes at a time) for ElasticSearch cluster search_codfw: apply logging and security config updates - bking@cumin2003 - [[phab:T324335|T324335]]
* 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 14:00 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 13:49 swfrench@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2330.codfw.wmnet
* 13:49 swfrench@cumin2002: conftool action : set/pooled=no; selector: name=wikikube-worker2330.codfw.wmnet
* 13:48 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] (duration: 09m 19s)
* 13:45 swfrench@dns1004: END - running authdns-update
* 13:44 mszwarc@deploy1003: mszwarc, jforrester: Continuing with deployment
* 13:43 swfrench@dns1004: START - running authdns-update
* 13:41 mszwarc@deploy1003: mszwarc, jforrester: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:38 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1320171{{!}}Thumbnail: Exclude map figures from multimediaviewer (T433703 T427709)]], [[gerrit:1320159{{!}}TimedText: Include SRT-only sources when listing playback tracks (T433666)]]
* 13:33 otto@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 13:33 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1154.eqiad.wmnet
* 13:32 otto@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 13:32 otto@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 13:31 otto@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 13:31 otto@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:31 otto@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:29 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1154.eqiad.wmnet
* 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1154.eqiad.wmnet
* 13:28 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1154.eqiad.wmnet
* 13:28 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1141.eqiad.wmnet
* 13:28 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1141.eqiad.wmnet
* 13:28 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1141.eqiad.wmnet
* 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 13:22 otto@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/mw-page-html-feature-counts-change-enrich: apply
* 13:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie
* 13:05 swfrench@dns1004: END - running authdns-update
* 13:03 swfrench@dns1004: START - running authdns-update
* 12:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 12:43 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1171.eqiad.wmnet with reason: decom
* 12:42 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 1:00:00 on db1150.eqiad.wmnet with reason: decom
* 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 12:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1164,1217].eqiad.wmnet with reason: cloning
* 12:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2096.codfw.wmnet with OS trixie
* 12:22 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie
* 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage
* 12:14 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1141.eqiad.wmnet with OS trixie
* 12:10 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2096.codfw.wmnet with reason: host reimage
* 12:10 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1289.eqiad.wmnet
* 12:05 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1289.eqiad.wmnet
* 12:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1288.eqiad.wmnet
* 11:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1288.eqiad.wmnet
* 11:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1287.eqiad.wmnet
* 11:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1097.eqiad.wmnet with OS trixie
* 11:54 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1287.eqiad.wmnet
* 11:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1286.eqiad.wmnet
* 11:53 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage
* 11:51 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2096.codfw.wmnet with OS trixie
* 11:49 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1141.eqiad.wmnet with reason: host reimage
* 11:48 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1286.eqiad.wmnet
* 11:48 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1284.eqiad.wmnet
* 11:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2095.codfw.wmnet with OS trixie
* 11:43 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1284.eqiad.wmnet
* 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1283.eqiad.wmnet
* 11:42 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on ml-serve1015.eqiad.wmnet with reason: Downtime to get full picture of current BIOS settings beyond what Redfish shows
* 11:39 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad
* 11:39 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:37 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1283.eqiad.wmnet
* 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1282.eqiad.wmnet
* 11:37 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad
* 11:37 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1141
* 11:33 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1141
* 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 11:32 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1141
* 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:32 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1141.eqiad.wmnet 156.48.64.10.in-addr.arpa 6.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:32 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003"
* 11:32 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1141 - jayme@cumin1003"
* 11:32 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1282.eqiad.wmnet
* 11:32 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1281.eqiad.wmnet
* 11:32 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.tls (exit_code=99) for network device lswtest-d8-eqiad
* 11:32 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 11:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage
* 11:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie
* 11:26 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1281.eqiad.wmnet
* 11:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1280.eqiad.wmnet
* 11:25 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1097.eqiad.wmnet with reason: host reimage
* 11:22 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1141
* 11:21 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1141.eqiad.wmnet with OS trixie
* 11:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1280.eqiad.wmnet
* 11:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1279.eqiad.wmnet
* 11:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin with reason: upgrade new Nokia swtiches in eqsin to SR Linux v26
* 11:17 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1141.eqiad.wmnet
* 11:16 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1141.eqiad.wmnet
* 11:16 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1141.eqiad.wmnet
* 11:16 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2095.codfw.wmnet with OS trixie
* 11:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1279.eqiad.wmnet
* 11:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1278.eqiad.wmnet
* 11:14 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1139.eqiad.wmnet
* 11:14 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1139.eqiad.wmnet
* 11:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1070.eqiad.wmnet with OS trixie
* 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1140.eqiad.wmnet
* 11:13 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1140.eqiad.wmnet
* 11:13 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1140.eqiad.wmnet
* 11:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1278.eqiad.wmnet
* 11:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1071.eqiad.wmnet with OS trixie
* 11:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie
* 11:04 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be1097.eqiad.wmnet with OS trixie
* 11:02 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1140.eqiad.wmnet with OS trixie
* 11:02 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1097.eqiad.wmnet with OS trixie
* 11:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1096.eqiad.wmnet with OS trixie
* 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1139.eqiad.wmnet
* 11:00 jayme@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1139.eqiad.wmnet with OS trixie
* 10:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1069.eqiad.wmnet with OS trixie
* 10:56 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage
* 10:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 10:45 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage
* 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1096.eqiad.wmnet with reason: host reimage
* 10:41 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1096.eqiad.wmnet with OS trixie
* 10:41 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage
* 10:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1071.eqiad.wmnet with reason: host reimage
* 10:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1070.eqiad.wmnet with reason: host reimage
* 10:38 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage
* 10:37 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1140.eqiad.wmnet with reason: host reimage
* 10:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage
* 10:33 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1095.eqiad.wmnet with OS trixie
* 10:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2095.codfw.wmnet with reason: host reimage
* 10:33 jayme@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1139.eqiad.wmnet with reason: host reimage
* 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1069.eqiad.wmnet with reason: host reimage
* 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1140
* 10:23 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1140
* 10:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 10:22 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1140
* 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1071.eqiad.wmnet with OS trixie
* 10:21 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1140.eqiad.wmnet 155.48.64.10.in-addr.arpa 5.5.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:21 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003"
* 10:21 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1140 - jayme@cumin1003"
* 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1071
* 10:21 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1070.eqiad.wmnet with OS trixie
* 10:21 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1070
* 10:20 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 10:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie
* 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1139
* 10:17 jayme@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1139
* 10:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be1095.eqiad.wmnet with OS trixie
* 10:15 jayme@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1139
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:15 jayme@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1139.eqiad.wmnet 194.32.64.10.in-addr.arpa 4.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:15 jayme@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003"
* 10:15 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 10:15 jayme@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1139 - jayme@cumin1003"
* 10:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2095.codfw.wmnet with OS trixie
* 10:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1069.eqiad.wmnet with OS trixie
* 10:12 mvernon@cumin1003: END (FAIL) - Cookbook sre.swift.convert-disks (exit_code=99) for host ms-be1069
* 10:11 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1140
* 10:11 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1140.eqiad.wmnet with OS trixie
* 10:11 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1140.eqiad.wmnet
* 10:10 jayme@cumin1003: START - Cookbook sre.dns.netbox
* 10:10 jayme@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1139
* 10:10 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1140.eqiad.wmnet
* 10:10 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1140.eqiad.wmnet
* 10:10 jayme@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1139.eqiad.wmnet with OS trixie
* 10:09 jayme@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1139.eqiad.wmnet
* 10:08 jayme@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1139.eqiad.wmnet
* 10:08 jayme@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1139.eqiad.wmnet
* 10:01 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2094.codfw.wmnet with OS trixie
* 09:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 09:54 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 09:53 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1095.eqiad.wmnet with reason: host reimage
* 09:53 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:44 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage
* 09:34 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2094.codfw.wmnet with reason: host reimage
* 09:34 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1071
* 09:33 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1070
* 09:33 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1095.eqiad.wmnet with OS trixie
* 09:27 mvernon@cumin1003: START - Cookbook sre.swift.convert-disks for host ms-be1069
* 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:24 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:23 brouberol@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM archiva1002.wikimedia.org
* 09:20 brouberol@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM archiva1002.wikimedia.org
* 09:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1277.eqiad.wmnet
* 09:13 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2094.codfw.wmnet with OS trixie
* 09:13 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 09:12 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:12 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1277.eqiad.wmnet
* 09:12 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 09:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1276.eqiad.wmnet
* 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1094.eqiad.wmnet with OS trixie
* 09:06 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1276.eqiad.wmnet
* 09:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1275.eqiad.wmnet
* 09:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2093.codfw.wmnet with OS trixie
* 09:01 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1275.eqiad.wmnet
* 09:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1274.eqiad.wmnet
* 08:56 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1274.eqiad.wmnet
* 08:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1273.eqiad.wmnet
* 08:50 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1273.eqiad.wmnet
* 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1272.eqiad.wmnet
* 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:49 dcausse@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage
* 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1272.eqiad.wmnet
* 08:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1094.eqiad.wmnet with reason: host reimage
* 08:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage
* 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:38 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2093.codfw.wmnet with reason: host reimage
* 08:35 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 08:34 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 08:29 dcausse@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:28 dcausse@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-semantic-search-test: apply
* 08:26 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1271.eqiad.wmnet
* 08:23 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1094.eqiad.wmnet with OS trixie
* 08:21 jnuche@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 08:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1271.eqiad.wmnet
* 08:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1270.eqiad.wmnet
* 08:15 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1270.eqiad.wmnet
* 08:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1269.eqiad.wmnet
* 08:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2093.codfw.wmnet with OS trixie
* 08:09 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1269.eqiad.wmnet
* 08:09 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1268.eqiad.wmnet
* 08:06 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2092.codfw.wmnet with OS trixie
* 08:04 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1268.eqiad.wmnet
* 08:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1267.eqiad.wmnet
* 07:59 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1267.eqiad.wmnet
* 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1093.eqiad.wmnet with OS trixie
* 07:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host db1266.eqiad.wmnet
* 07:56 jynus: running extra backups to test db1285 [[phab:T433826|T433826]]
* 07:51 cwilliams@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1266.eqiad.wmnet
* 07:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage
* 07:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage
* 07:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2092.codfw.wmnet with reason: host reimage
* 07:32 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1093.eqiad.wmnet with reason: host reimage
* 07:29 jynus: running extra backups to test db1265 [[phab:T433825|T433825]]
* 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2092.codfw.wmnet with OS trixie
* 07:11 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1093.eqiad.wmnet with OS trixie
* 06:50 slyngshede@dns1004: END - running authdns-update
* 06:48 slyngshede@dns1004: START - running authdns-update
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.11 (duration: 02m 29s)
* 03:38 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]] (duration: 32m 57s)
* 03:23 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 03:22 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 03:05 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.14 refs [[phab:T430833|T430833]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 32s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:45 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] (duration: 06m 20s)
* 00:41 cjming@deploy1003: cjming: Continuing with deployment
* 00:41 cjming@deploy1003: cjming: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 00:39 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320296{{!}}Fix InstrumentConstructiveEdits script (T431493)]]
== 2026-08-03 ==
* 23:58 rzl@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-cron: apply
* 23:57 rzl@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-cron: apply
* 23:29 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:29 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:27 robh@cumin1003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:26 robh@cumin1003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts cp5021.eqsin.wmnet
* 23:18 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 23:17 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 22:56 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: sync
* 22:56 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: sync
* 22:36 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 22:36 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 22:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host search-loader2002.codfw.wmnet with OS trixie
* 21:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage
* 21:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on search-loader2002.codfw.wmnet with reason: host reimage
* 21:42 dancy@deploy1003: Stopping before sync operations
* 21:41 dancy@deploy1003: Started scap sync-world: testing
* 21:39 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts
* 21:37 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s)
* 21:37 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] (duration: 06m 13s)
* 21:33 dancy@deploy1003: dancy: Continuing with deployment
* 21:32 dancy@deploy1003: dancy: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:31 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1320233{{!}}mergeMessageFileList: Suppress "doesn't exist" error when --quiet (T125678 T433104)]]
* 21:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host search-loader2002.codfw.wmnet with OS trixie
* 21:03 dancy@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] (duration: 06m 34s)
* 20:59 dancy@deploy1003: dancy: Continuing with deployment
* 20:58 dancy@deploy1003: dancy: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:56 dancy@deploy1003: Started scap sync-world: Backport for [[gerrit:1315968{{!}}wmf-config/logging.php: Adjustments for WMF_MAINTENANCE_OFFLINE (T433104)]]
* 20:52 cjming@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] (duration: 06m 23s)
* 20:48 cjming@deploy1003: cjming: Continuing with deployment
* 20:47 cjming@deploy1003: cjming: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:46 cjming@deploy1003: Started scap sync-world: Backport for [[gerrit:1320215{{!}}eventLogUtils: Set schema on Test Kitchen instrument before send (T433550)]]
* 20:42 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] (duration: 07m 36s)
* 20:38 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:36 arlolra@deploy1003: arlolra: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:34 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1320179{{!}}Turn on PRV for all namespaces on enwiki (T430194)]]
* 20:16 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] (duration: 08m 26s)
* 20:12 krinkle@deploy1003: krinkle: Continuing with deployment
* 20:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319903{{!}}Re-enable wgTrackMediaRequestProvenance on Commons and Wikipedia (T414338)]]
* 19:45 jasmine@cumin2002: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-main-codfw
* 18:58 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] (duration: 09m 23s)
* 18:53 krinkle@deploy1003: krinkle: Continuing with deployment
* 18:53 jasmine@cumin2002: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-main-codfw
* 18:50 krinkle@deploy1003: krinkle: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:48 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1311041{{!}}MathML+MathJax rollout to phase 2 (not Wikibooks/Wikisource/Wikipedia) (T271001)]]
* 18:37 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] (duration: 10m 13s)
* 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch2001.codfw.wmnet
* 18:34 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch2001.codfw.wmnet with OS trixie
* 18:33 krinkle@deploy1003: krinkle: Continuing with deployment
* 18:29 krinkle@deploy1003: krinkle: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:27 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1319868{{!}}logging: Remove mention of 'fatal' channel that no longer exists (T247113)]]
* 18:19 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 18:18 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage
* 18:14 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 18:14 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:12 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch2001.codfw.wmnet with reason: host reimage
* 18:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 18:11 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 18:10 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 18:02 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 18:02 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 18:01 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 18:01 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 17:55 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch2001.codfw.wmnet with OS trixie
* 17:54 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:54 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch2001.codfw.wmnet on all recursors
* 17:53 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch2001.codfw.wmnet on all recursors
* 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:53 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:48 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch2001.codfw.wmnet - dzahn@cumin2002"
* 17:41 dzahn@cumin2002: START - Cookbook sre.dns.netbox
* 17:41 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch2001.codfw.wmnet
* 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.ganeti.makevm (exit_code=0) for new host codesearch1001.eqiad.wmnet
* 17:37 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host codesearch1001.eqiad.wmnet with OS trixie
* 17:24 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage
* 17:17 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on codesearch1001.eqiad.wmnet with reason: host reimage
* 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host codesearch1001.eqiad.wmnet with OS trixie
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 17:06 rzl@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:06 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.ganeti.makevm: created new VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) codesearch1001.eqiad.wmnet on all recursors
* 17:06 dzahn@cumin2002: START - Cookbook sre.dns.wipe-cache codesearch1001.eqiad.wmnet on all recursors
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:06 dzahn@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 17:05 rzl@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 17:04 rzl@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:04 rzl@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 16:58 dzahn@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add records for VM codesearch1001.eqiad.wmnet - dzahn@cumin2002"
* 16:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2091.codfw.wmnet with OS trixie
* 16:54 ebernhardson@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 16:54 ebernhardson@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 16:49 ebernhardson@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 16:49 ebernhardson@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 16:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1092.eqiad.wmnet with OS trixie
* 16:43 ebernhardson@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 16:43 ebernhardson@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 16:43 dzahn@cumin2002: START - Cookbook sre.dns.netbox
* 16:43 dzahn@cumin2002: START - Cookbook sre.ganeti.makevm for new host codesearch1001.eqiad.wmnet
* 16:41 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc2046
* 16:41 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc2046
* 16:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage
* 16:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 16:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2091.codfw.wmnet with reason: host reimage
* 16:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage
* 16:24 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 16:24 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 16:23 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1092.eqiad.wmnet with reason: host reimage
* 16:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2091.codfw.wmnet with OS trixie
* 16:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1092.eqiad.wmnet with OS trixie
* 16:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2090.codfw.wmnet with OS trixie
* 15:51 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 15:51 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 15:51 jiji@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 15:50 jiji@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage
* 15:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2090.codfw.wmnet with reason: host reimage
* 15:33 jhathaway@dns1004: END - running authdns-update
* 15:31 jhathaway@dns1004: START - running authdns-update
* 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1091.eqiad.wmnet with OS trixie
* 15:25 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts
* 15:23 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s)
* 15:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2090.codfw.wmnet with OS trixie
* 15:12 marostegui@cumin1003: dbctl commit (dc=all): 'Repool db2245, db2246, db2247 and db2248 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95857 and previous config saved to /var/cache/conftool/dbconfig/20260803-151212-marostegui.json
* 15:09 dancy@deploy1003: Started scap sync-world: testing
* 15:09 dancy@deploy1003: Installation of scap version "4.277.0" completed for 3 hosts
* 15:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage
* 15:07 dancy@deploy1003: Installing scap version "4.277.0" for 3 host(s)
* 15:03 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1091.eqiad.wmnet with reason: host reimage
* 14:49 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1091.eqiad.wmnet with OS trixie
* 14:33 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2089.codfw.wmnet with OS trixie
* 14:29 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 14:27 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 14:18 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 14:16 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 14:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage
* 14:10 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 14:10 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 14:09 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2089.codfw.wmnet with reason: host reimage
* 13:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2089.codfw.wmnet with OS trixie
* 13:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2088.codfw.wmnet with OS trixie
* 13:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1090.eqiad.wmnet with OS trixie
* 13:22 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage
* 13:22 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] (duration: 14m 34s)
* 13:19 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1090.eqiad.wmnet with reason: host reimage
* 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage
* 13:16 aude@deploy1003: aude, mhorsey: Continuing with deployment
* 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2088.codfw.wmnet with reason: host reimage
* 13:12 aude@deploy1003: aude, mhorsey: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:08 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1320149{{!}}Enable Special:ChartWizard on Wikimedia Commons (T433831)]], [[gerrit:1319854{{!}}enable CampaignEvents worklists (T429507 T429508)]]
* 13:05 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1090.eqiad.wmnet with OS trixie
* 12:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2088.codfw.wmnet with OS trixie
* 12:56 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db[2245-2247].codfw.wmnet
* 12:49 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2247: Rebooting db2247.codfw.wmnet
* 12:49 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2247: Rebooting db2247.codfw.wmnet
* 12:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2246: Rebooting db2246.codfw.wmnet
* 12:42 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2246: Rebooting db2246.codfw.wmnet
* 12:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2087.codfw.wmnet with OS trixie
* 12:37 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1089.eqiad.wmnet with OS trixie
* 12:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2245: Rebooting db2245.codfw.wmnet
* 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2245: Rebooting db2245.codfw.wmnet
* 12:34 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db[2245-2247].codfw.wmnet
* 12:32 kamila@deploy1003: Finished scap sync-world: rebuild after base image update (duration: 30m 26s)
* 12:28 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 12:22 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage
* 12:19 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage
* 12:14 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2087.codfw.wmnet with reason: host reimage
* 12:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1089.eqiad.wmnet with reason: host reimage
* 12:03 kamila@deploy1003: Started scap sync-world: rebuild after base image update
* 12:00 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1089.eqiad.wmnet with OS trixie
* 12:00 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2087.codfw.wmnet with OS trixie
* 11:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2245-2248].codfw.wmnet
* 11:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2245-2248].codfw.wmnet
* 11:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2086.codfw.wmnet with OS trixie
* 11:26 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:26 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 11:25 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:25 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 11:24 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1088.eqiad.wmnet with OS trixie
* 11:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db[2245-2248].codfw.wmnet with reason: Checking network
* 11:21 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:20 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 11:19 marostegui@dns1004: END - running authdns-update
* 11:17 marostegui@dns1004: START - running authdns-update
* 11:10 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:10 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 11:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage
* 11:09 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:08 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop-jobqueue: apply
* 11:08 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 11:07 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop-jobqueue: apply
* 11:07 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop-jobqueue: apply
* 11:06 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop: apply
* 11:06 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop: apply
* 11:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage
* 11:05 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/changeprop: apply
* 11:04 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/changeprop: apply
* 11:04 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/changeprop: apply
* 11:04 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/changeprop: apply
* 11:02 marostegui@dns1004: END - running authdns-update
* 11:02 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2086.codfw.wmnet with reason: host reimage
* 11:01 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1088.eqiad.wmnet with reason: host reimage
* 11:00 marostegui@dns1004: START - running authdns-update
* 10:53 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] (duration: 10m 57s)
* 10:51 cmooney@dns3003: END - running authdns-update
* 10:49 cmooney@dns3003: START - running authdns-update
* 10:47 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1088.eqiad.wmnet with OS trixie
* 10:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2086.codfw.wmnet with OS trixie
* 10:47 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 10:46 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003"
* 10:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new reverse ranges for eqsin CR switch links - cmooney@cumin1003"
* 10:42 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1320115{{!}}Remove more $wmg = $wg hacks (T119117)]]
* 10:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 10:36 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2245, db2246 and db2247 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95855 and previous config saved to /var/cache/conftool/dbconfig/20260803-103652-marostegui.json
* 10:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2248 from s4 [[phab:T433610|T433610]]', diff saved to https://phabricator.wikimedia.org/P95854 and previous config saved to /var/cache/conftool/dbconfig/20260803-103535-marostegui.json
* 10:27 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/rest-gateway: apply
* 10:27 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/rest-gateway: apply
* 10:26 cgoubert@deploy1003: helmfile [codfw] DONE helmfile.d/services/rest-gateway: apply
* 10:24 kart_: cxserver: Add referencePunctuation config ([[phab:T97231|T97231]])
* 10:24 cgoubert@deploy1003: helmfile [codfw] START helmfile.d/services/rest-gateway: apply
* 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 10:23 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 10:23 cgoubert@deploy1003: helmfile [staging] DONE helmfile.d/services/rest-gateway: apply
* 10:22 cgoubert@deploy1003: helmfile [staging] START helmfile.d/services/rest-gateway: apply
* 10:22 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 10:21 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2085.codfw.wmnet with OS trixie
* 10:20 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 10:20 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 10:18 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 10:18 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 10:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1087.eqiad.wmnet with OS trixie
* 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage
* 09:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage
* 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2085.codfw.wmnet with reason: host reimage
* 09:40 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1087.eqiad.wmnet with reason: host reimage
* 09:26 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1087.eqiad.wmnet with OS trixie
* 09:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2085.codfw.wmnet with OS trixie
* 09:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2084.codfw.wmnet with OS trixie
* 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1086.eqiad.wmnet with OS trixie
* 08:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage
* 08:50 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2084.codfw.wmnet with reason: host reimage
* 08:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage
* 08:39 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1086.eqiad.wmnet with reason: host reimage
* 08:38 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:38 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:37 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:37 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:35 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2084.codfw.wmnet with OS trixie
* 08:34 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:34 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1086.eqiad.wmnet with OS trixie
* 08:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Repool after a crash
* 08:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2083.codfw.wmnet with OS trixie
* 08:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1085.eqiad.wmnet with OS trixie
* 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage
* 07:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage
* 07:40 kart_: Updated cxsever to 2026-07-16-140518-production ([[phab:T97231|T97231]])
* 07:39 kartik@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cxserver: apply
* 07:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2083.codfw.wmnet with reason: host reimage
* 07:38 kartik@deploy1003: helmfile [eqiad] START helmfile.d/services/cxserver: apply
* 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1085.eqiad.wmnet with reason: host reimage
* 07:37 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] (duration: 32m 40s)
* 07:33 kartik@deploy1003: helmfile [codfw] DONE helmfile.d/services/cxserver: apply
* 07:33 kartik@deploy1003: helmfile [codfw] START helmfile.d/services/cxserver: apply
* 07:25 jdlrobson@deploy1003: jdlrobson: Continuing with deployment
* 07:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2083.codfw.wmnet with OS trixie
* 07:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1085.eqiad.wmnet with OS trixie
* 07:23 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Repool after a crash
* 07:21 jdlrobson@deploy1003: jdlrobson: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:09 marostegui: Drop renamed tables [[phab:T425074|T425074]]
* 07:04 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1315895{{!}}Enable page images on recipe namespace]]
* 06:55 kartik@deploy1003: helmfile [staging] DONE helmfile.d/services/cxserver: apply
* 06:54 kartik@deploy1003: helmfile [staging] START helmfile.d/services/cxserver: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 46s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-02 ==
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 03s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-08-01 ==
* 03:30 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:30 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:30 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:30 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 34s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-31 ==
* 17:41 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 17:41 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 17:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 17:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 15:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Testing
* 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:02 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 15:02 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 14:48 pt1979@cumin2002: START - Cookbook sre.dns.netbox
* 14:47 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Testing
* 14:22 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 14:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing
* 14:21 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 14:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2195.codfw.wmnet with reason: Testing
* 14:16 aikochou@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 14:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 14:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 13:30 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cloudvirt1048.eqiad.wmnet with OS trixie
* 13:22 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing
* 13:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 13:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db2195: Testing
* 13:18 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 13:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2195: Testing
* 13:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2195: Testing
* 13:05 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 13:05 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 13:04 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 13:04 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 12:50 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:49 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lswtest-d8-eqiad
* 12:49 ayounsi@cumin1003: START - Cookbook sre.network.tls for network device lswtest-d8-eqiad
* 11:53 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:42 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 6515
* 11:37 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 6515
* 11:28 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:27 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:07 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 11:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2082.codfw.wmnet with OS trixie
* 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage
* 10:42 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2082.codfw.wmnet with reason: host reimage
* 10:28 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2082.codfw.wmnet with OS trixie
* 10:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 09:52 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 09:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts
* 09:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts
* 08:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 08:46 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:42 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revise-tone-task-generator' for release 'main' .
* 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 08:37 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 08:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 08:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 08:11 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:11 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 08:08 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048
* 08:07 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 08:07 filippo@cumin1003: END (FAIL) - Cookbook sre.network.configure-switch-interfaces (exit_code=99) for host cloudvirt1048
* 08:06 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 08:01 filippo@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 08:00 filippo@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 07:19 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage
* 07:13 filippo@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cloudvirt1048.eqiad.wmnet with reason: host reimage
* 07:11 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 07:11 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 07:09 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 07:09 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "cloudvirt1048 - filippo@cumin1003"
* 06:57 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:56 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:48 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 06:44 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie
* 06:34 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 00:57 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] (duration: 11m 04s)
* 00:53 dreamyjazz@deploy1003: dreamyjazz, jforrester: Continuing with deployment
* 00:48 dreamyjazz@deploy1003: dreamyjazz, jforrester: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 00:46 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1316832{{!}}Drop old support for non-temporary accounts wikis]]
== 2026-07-30 ==
* 21:37 dancy@deploy1003: Installation of scap version "4.276.1" completed for 3 hosts
* 21:35 dancy@deploy1003: Installing scap version "4.276.1" for 3 host(s)
* 21:24 dancy@deploy1003: Installation of scap version "4.276.0" completed for 3 hosts
* 21:22 dancy@deploy1003: Installing scap version "4.276.0" for 3 host(s)
* 21:15 maryum: Deployed security fix for [[phab:T430601|T430601]]
* 20:13 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] (duration: 09m 20s)
* 20:07 arlolra@deploy1003: osleger, arlolra: Continuing with deployment
* 20:05 arlolra@deploy1003: osleger, arlolra: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:03 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1319521{{!}}Increase Parsoid image limit to 5000 (T430854)]]
* 19:29 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:29 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:25 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service
* 19:24 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 19:24 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:24 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host ms-be2082.mgmt.codfw.wmnet with chassis set policy GRACEFUL_RESTART
* 19:24 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 19:23 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 19:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1084.eqiad.wmnet with OS trixie
* 18:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage
* 18:52 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1084.eqiad.wmnet with reason: host reimage
* 18:41 cjming@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:40 cjming@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 18:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1084.eqiad.wmnet with OS trixie
* 18:25 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 18:15 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 17:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1083.eqiad.wmnet with OS trixie
* 17:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1048: Maintenance
* 17:36 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]]
* 17:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage
* 17:26 inflatador: bking@apt1002 `reprepro --noskipold --component thirdparty/opensearch3 update trixie-wikimedia` [[phab:T433624|T433624]]
* 17:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1083.eqiad.wmnet with reason: host reimage
* 17:23 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 17:20 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 17:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2081.codfw.wmnet with OS trixie
* 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new security plugin settings - bking@cumin2003 - [[phab:T350516|T350516]]
* 17:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1083.eqiad.wmnet with OS trixie
* 16:55 root@cumin1003: START - Cookbook sre.mysql.pool pool es1048: Maintenance
* 16:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage
* 16:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95833 and previous config saved to /var/cache/conftool/dbconfig/20260730-165053-cwilliams.json
* 16:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1048.eqiad.wmnet with reason: Maintenance
* 16:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1040: Maintenance
* 16:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2081.codfw.wmnet with reason: host reimage
* 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1082.eqiad.wmnet with OS trixie
* 16:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2081.codfw.wmnet with OS trixie
* 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2097.codfw.wmnet with OS trixie
* 16:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage
* 16:08 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1082.eqiad.wmnet with reason: host reimage
* 16:08 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 16:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1047: Maintenance
* 16:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2080.codfw.wmnet with OS trixie
* 16:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance
* 16:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1040: Maintenance
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:58 root@cumin1003: START - Cookbook sre.mysql.pool pool es1040: Maintenance
* 15:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95827 and previous config saved to /var/cache/conftool/dbconfig/20260730-155324-cwilliams.json
* 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1040.eqiad.wmnet with reason: Maintenance
* 15:50 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1082.eqiad.wmnet with OS trixie
* 15:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2048: Maintenance
* 15:44 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s)
* 15:43 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage
* 15:38 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.RESTART (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply new logging settings - bking@cumin2003 - [[phab:T324335|T324335]]
* 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2080.codfw.wmnet with reason: host reimage
* 15:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 15:30 mvernon@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS trixie
* 15:23 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1081.eqiad.wmnet with OS trixie
* 15:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2098.codfw.wmnet with OS trixie
* 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003"
* 15:18 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie
* 15:18 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - mvernon@cumin2003"
* 15:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 15:17 root@cumin1003: START - Cookbook sre.mysql.pool pool es1047: Maintenance
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2080.codfw.wmnet with OS trixie
* 15:13 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be2097.codfw.wmnet with OS trixie
* 15:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95820 and previous config saved to /var/cache/conftool/dbconfig/20260730-151200-cwilliams.json
* 15:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1047.eqiad.wmnet with reason: Maintenance
* 15:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: Maintenance
* 15:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage
* 15:00 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1081.eqiad.wmnet with reason: host reimage
* 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool es2048: Maintenance
* 15:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 14:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2079.codfw.wmnet with OS trixie
* 14:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2048 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95816 and previous config saved to /var/cache/conftool/dbconfig/20260730-145510-cwilliams.json
* 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2048.codfw.wmnet with reason: Maintenance
* 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2040: Maintenance
* 14:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 14:51 tchin@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] (duration: 06m 48s)
* 14:47 tchin@deploy1003: jforrester, tchin: Continuing with deployment
* 14:47 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 14:47 tchin@deploy1003: jforrester, tchin: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:45 tchin@deploy1003: Started scap sync-world: Backport for [[gerrit:1319489{{!}}Monolog: Strip diagnostic context from Monolog-based EventBus events (T433457)]]
* 14:42 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid
* 14:42 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid
* 14:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1081.eqiad.wmnet with OS trixie
* 14:39 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 14:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage
* 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS trixie
* 14:32 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2079.codfw.wmnet with reason: host reimage
* 14:30 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] (duration: 06m 31s)
* 14:27 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS trixie
* 14:26 mszwarc@deploy1003: mszwarc: Continuing with deployment
* 14:25 mszwarc@deploy1003: mszwarc: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:25 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance
* 14:25 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1038: Maintenance
* 14:23 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1319482{{!}}plwikiquote: Set AutoConfirmCount to 25 (T433541)]], [[gerrit:1309552{{!}}Revert "Temporarily change plwiki tagline for 1.7M articles" (T430512)]]
* 14:21 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] (duration: 11m 19s)
* 14:20 root@cumin1003: START - Cookbook sre.mysql.pool pool es1038: Maintenance
* 14:14 stran@deploy1003: stran: Continuing with deployment
* 14:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95810 and previous config saved to /var/cache/conftool/dbconfig/20260730-141439-cwilliams.json
* 14:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1038.eqiad.wmnet with reason: Maintenance
* 14:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1036: Maintenance
* 14:13 stran@deploy1003: stran: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2079.codfw.wmnet with OS trixie
* 14:09 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319487{{!}}SI: Add performer to server-side link_click event (T433257)]], [[gerrit:1319488{{!}}SI: Add performer to server-side link_click event (T433257)]]
* 14:08 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance
* 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2040: Maintenance
* 14:03 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] (duration: 31m 41s)
* 14:03 root@cumin1003: START - Cookbook sre.mysql.pool pool es2040: Maintenance
* 14:03 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2040 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95806 and previous config saved to /var/cache/conftool/dbconfig/20260730-135643-cwilliams.json
* 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2040.codfw.wmnet with reason: Maintenance
* 13:56 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:56 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2038: Maintenance
* 13:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:52 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Continuing with deployment
* 13:49 lucaswerkmeister-wmde@deploy1003: migr, lucaswerkmeister-wmde: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:49 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:48 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:45 kevinbazira@deploy1003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:32 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:32 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319449{{!}}SpecialCreateAccount: make user policy-link available again (T430604)]]
* 13:28 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance
* 13:28 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es1036: Maintenance
* 13:22 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:22 root@cumin1003: START - Cookbook sre.mysql.pool pool es1036: Maintenance
* 13:20 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es1036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95800 and previous config saved to /var/cache/conftool/dbconfig/20260730-131727-cwilliams.json
* 13:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1036.eqiad.wmnet with reason: Maintenance
* 13:17 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] (duration: 10m 31s)
* 13:16 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2022\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 13:13 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Continuing with deployment
* 13:10 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T433606|T433606]]
* 13:10 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance
* 13:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2038: Maintenance
* 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, stran: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie
* 13:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2047: Maintenance
* 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003"
* 13:07 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048 cloud-private - filippo@cumin1003"
* 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1319466{{!}}Instrument link_click server-side instead of client-side (T433257)]], [[gerrit:1319467{{!}}Instrument link_click server-side instead of client-side (T433257)]]
* 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool es2038: Maintenance
* 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2078.codfw.wmnet with OS trixie
* 13:01 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2038 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95797 and previous config saved to /var/cache/conftool/dbconfig/20260730-125919-cwilliams.json
* 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2038.codfw.wmnet with reason: Maintenance
* 12:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage
* 12:37 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2078.codfw.wmnet with reason: host reimage
* 12:37 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] (duration: 06m 51s)
* 12:33 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:32 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 12:32 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:32 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:30 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1319162{{!}}Remove $wmg hack for UploadStashMaxAge (T119117)]]
* 12:19 root@cumin1003: START - Cookbook sre.mysql.pool pool es2047: Maintenance
* 12:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2078.codfw.wmnet with OS trixie
* 12:18 dcausse@deploy1003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 12:18 dcausse@deploy1003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 12:15 dcausse@deploy1003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 12:14 dcausse@deploy1003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 12:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2047 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95793 and previous config saved to /var/cache/conftool/dbconfig/20260730-121404-cwilliams.json
* 12:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2047.codfw.wmnet with reason: Maintenance
* 12:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2036: Maintenance
* 12:05 ayounsi@dns1004: END - running authdns-update
* 12:02 ayounsi@dns1004: START - running authdns-update
* 11:51 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2077.codfw.wmnet with OS trixie
* 11:48 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:46 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1080.eqiad.wmnet with OS trixie
* 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1226: Maintenance
* 11:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage
* 11:28 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2077.codfw.wmnet with reason: host reimage
* 11:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage
* 11:27 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance
* 11:27 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2036: Maintenance
* 11:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1080.eqiad.wmnet with reason: host reimage
* 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool es2036: Maintenance
* 11:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2036 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95786 and previous config saved to /var/cache/conftool/dbconfig/20260730-111633-cwilliams.json
* 11:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2036.codfw.wmnet with reason: Maintenance
* 11:08 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2077.codfw.wmnet with OS trixie
* 11:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 11:03 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie
* 11:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1226: Maintenance
* 10:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95783 and previous config saved to /var/cache/conftool/dbconfig/20260730-104801-cwilliams.json
* 10:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1226.eqiad.wmnet with reason: Maintenance
* 10:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1214: Maintenance
* 10:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: Maintenance
* 10:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2076.codfw.wmnet with OS trixie
* 10:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1214: Maintenance
* 09:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1214 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95775 and previous config saved to /var/cache/conftool/dbconfig/20260730-095451-cwilliams.json
* 09:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1214.eqiad.wmnet with reason: Maintenance
* 09:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1209: Maintenance
* 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage
* 09:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance
* 09:42 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox
* 09:41 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox
* 09:40 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool es2035: Maintenance
* 09:39 ayounsi@cumin1003: END (PASS) - Cookbook sre.netbox.update-extras (exit_code=0) rolling restart_daemons on A:netbox-canary
* 09:39 ayounsi@cumin1003: START - Cookbook sre.netbox.update-extras rolling restart_daemons on A:netbox-canary
* 09:39 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2076.codfw.wmnet with reason: host reimage
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:35 lucaswerkmeister-wmde@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 09:34 root@cumin1003: START - Cookbook sre.mysql.pool pool es2035: Maintenance
* 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 09:32 lucaswerkmeister-wmde@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 09:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling es2035 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95771 and previous config saved to /var/cache/conftool/dbconfig/20260730-092910-cwilliams.json
* 09:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es2035.codfw.wmnet with reason: Maintenance
* 09:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2076.codfw.wmnet with OS trixie
* 09:18 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 09:17 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie
* 09:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db1209: Maintenance
* 09:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 23 hosts
* 09:04 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003
* 09:04 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:02 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Remove cable label from interfaces descriptions - ayounsi@cumin1003
* 09:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1209 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95767 and previous config saved to /var/cache/conftool/dbconfig/20260730-090133-cwilliams.json
* 09:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1209.eqiad.wmnet with reason: Maintenance
* 09:01 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1192: Maintenance
* 08:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1252: Maintenance
* 08:57 jayme@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 08:56 jayme@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 08:53 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 23 hosts
* 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:51 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1263: Maintenance
* 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 08:23 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 08:17 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 08:15 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host ms-be1080.eqiad.wmnet with OS trixie
* 08:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db1192: Maintenance
* 08:13 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2075.codfw.wmnet with OS trixie
* 08:12 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1252: Maintenance
* 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1252: Maintenance
* 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1252: Maintenance
* 08:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1192 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95754 and previous config saved to /var/cache/conftool/dbconfig/20260730-080611-cwilliams.json
* 08:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1192.eqiad.wmnet with reason: Maintenance
* 08:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1178: Maintenance
* 08:05 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS bullseye
* 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1263: Maintenance
* 07:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1263 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95751 and previous config saved to /var/cache/conftool/dbconfig/20260730-075106-cwilliams.json
* 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[1260-1262].eqiad.wmnet with reason: Maintenance
* 07:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage
* 07:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1263.eqiad.wmnet with reason: Maintenance
* 07:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2075.codfw.wmnet with reason: host reimage
* 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance
* 07:38 dcausse@deploy1003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:38 dcausse@deploy1003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:35 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1252', diff saved to https://phabricator.wikimedia.org/P95748 and previous config saved to /var/cache/conftool/dbconfig/20260730-073510-marostegui.json
* 07:26 klausman@dns2004: END - running authdns-update
* 07:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1080.eqiad.wmnet with OS trixie
* 07:25 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2075.codfw.wmnet with OS trixie
* 07:24 klausman@dns2004: START - running authdns-update
* 07:23 btullis@cumin1003: END (ERROR) - Cookbook sre.hosts.reboot-single (exit_code=97) for host an-test-master1003.eqiad.wmnet
* 07:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1178: Maintenance
* 07:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95746 and previous config saved to /var/cache/conftool/dbconfig/20260730-071112-cwilliams.json
* 07:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1178.eqiad.wmnet with reason: Maintenance
* 07:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1177: Maintenance
* 06:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1177: Maintenance
* 06:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95741 and previous config saved to /var/cache/conftool/dbconfig/20260730-061736-cwilliams.json
* 06:17 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1177.eqiad.wmnet with reason: Maintenance
* 06:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1172: Maintenance
* 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: crashed
* 05:41 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db1217 it crashed', diff saved to https://phabricator.wikimedia.org/P95737 and previous config saved to /var/cache/conftool/dbconfig/20260730-054111-marostegui.json
* 05:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95736 and previous config saved to /var/cache/conftool/dbconfig/20260730-053422-cwilliams.json
* 05:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1172: Maintenance
* 05:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95734 and previous config saved to /var/cache/conftool/dbconfig/20260730-052414-cwilliams.json
* 05:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95733 and previous config saved to /var/cache/conftool/dbconfig/20260730-052354-cwilliams.json
* 05:23 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1172.eqiad.wmnet with reason: Maintenance
* 05:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1167: Maintenance
* 05:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95731 and previous config saved to /var/cache/conftool/dbconfig/20260730-051406-cwilliams.json
* 05:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95729 and previous config saved to /var/cache/conftool/dbconfig/20260730-050358-cwilliams.json
* 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 04:47 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 04:47 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add mr1 ge-0/0/3 ipv4 - pt1979@cumin2002"
* 04:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1167: Maintenance
* 04:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95726 and previous config saved to /var/cache/conftool/dbconfig/20260730-042923-cwilliams.json
* 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 04:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1167.eqiad.wmnet with reason: Maintenance
* 04:22 pt1979@cumin2002: START - Cookbook sre.dns.netbox
* 04:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95725 and previous config saved to /var/cache/conftool/dbconfig/20260730-040337-cwilliams.json
* 04:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:38 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]]
* 01:38 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: pool eqsin [reason: Switch upgrade maintenance window complete, [[phab:T433097|T433097]]]
== 2026-07-29 ==
* 23:57 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin,mr1-eqsin IPv6,mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: connection issue
* 22:54 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:53 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:53 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:53 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:25 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards
* 22:20 brett@cumin2002: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]]
* 22:20 brett@cumin2002: START - Cookbook sre.dns.admin DNS admin: depool eqsin [reason: Switch upgrade maintenance window, [[phab:T433097|T433097]]]
* 22:01 apine@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 22:00 apine@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 21:59 apine@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 21:58 apine@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 21:58 apine@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 21:58 apine@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 21:32 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host cloudvirt1048.eqiad.wmnet with OS trixie
* 21:25 pt1979@cumin2002: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host ms-be2097.codfw.wmnet with OS bullseye
* 21:16 zabe@deploy1003: mwscript-k8s job started: extensions/Translate/scripts/moveTranslatableBundle.php --wiki=metawiki 'Mental Health Resource Center' 'Safety Resource Center/Mental Health' Zabe --reason 'per request [[:phab:T433118{{!}}T433118]]'
* 21:12 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2022.codfw.wmnet, repooling source-only afterwards
* 21:12 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] (duration: 12m 53s)
* 21:12 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 21:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1253: Maintenance
* 21:08 aaron@deploy1003: aaron: Continuing with deployment
* 21:01 aaron@deploy1003: aaron: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:59 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1319175{{!}}Make the "wikibase-rest/v1" external module published (T422405)]]
* 20:52 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] (duration: 21m 57s)
* 20:48 aaron@deploy1003: aaron: Continuing with deployment
* 20:32 aaron@deploy1003: aaron: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:30 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1312694{{!}}Use the openapi.json endpoint for the wikibase-rest/v1 REST module (T422405)]]
* 20:24 root@cumin1003: START - Cookbook sre.mysql.pool pool db1253: Maintenance
* 20:19 aaron@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] (duration: 08m 07s)
* 20:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1253 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95719 and previous config saved to /var/cache/conftool/dbconfig/20260729-201810-cwilliams.json
* 20:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1253.eqiad.wmnet with reason: Maintenance
* 20:17 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1231: Maintenance
* 20:15 aaron@deploy1003: bpirkle, aaron: Continuing with deployment
* 20:13 aaron@deploy1003: bpirkle, aaron: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:12 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host cloudvirt1048.eqiad.wmnet with OS trixie
* 20:11 aaron@deploy1003: Started scap sync-world: Backport for [[gerrit:1306988{{!}}REST: use RestExternalModules config variable (T433314 T428375)]]
* 20:11 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:09 vriley@cumin1003: START - Cookbook sre.hosts.provision for host cloudvirt1048.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 20:09 pt1979@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 20:04 pt1979@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 19:47 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 19:43 pt1979@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye
* 19:41 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 19:37 zabe: zabe@deploy1003:~$ mwscript-k8s --comment='[[phab:T433529|T433529]]' --follow -- resetAuthenticationThrottle.php --wiki=aawiki --signup --ip=89.36.114.94
* 19:36 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] (duration: 06m 49s)
* 19:32 zabe@deploy1003: zabe: Continuing with deployment
* 19:31 zabe@deploy1003: zabe: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:31 root@cumin1003: START - Cookbook sre.mysql.pool pool db1231: Maintenance
* 19:29 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1319159{{!}}Add throttle exemption for Black Cultural Archives UK (T433529)]]
* 19:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95714 and previous config saved to /var/cache/conftool/dbconfig/20260729-192454-cwilliams.json
* 19:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1231.eqiad.wmnet with reason: Maintenance
* 19:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1227: Maintenance
* 19:22 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cloudvirt1048
* 19:22 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cloudvirt1048
* 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:21 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003"
* 19:21 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [cloudvirt1048] - vriley@cumin1003"
* 19:19 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 19:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1251: Maintenance
* 19:17 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95711 and previous config saved to /var/cache/conftool/dbconfig/20260729-191756-cwilliams.json
* 19:16 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:11 dduvall: rolling back wmf.13 to group0 due to [[phab:T433457|T433457]] (cc [[phab:T430832|T430832]])
* 19:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95709 and previous config saved to /var/cache/conftool/dbconfig/20260729-190748-cwilliams.json
* 19:01 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2022.codfw.wmnet with OS bookworm
* 19:01 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards
* 18:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252', diff saved to https://phabricator.wikimedia.org/P95707 and previous config saved to /var/cache/conftool/dbconfig/20260729-185740-cwilliams.json
* 18:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95704 and previous config saved to /var/cache/conftool/dbconfig/20260729-184732-cwilliams.json
* 18:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1227: Maintenance
* 18:34 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage
* 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95701 and previous config saved to /var/cache/conftool/dbconfig/20260729-183117-cwilliams.json
* 18:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1227.eqiad.wmnet with reason: Maintenance
* 18:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1202: Maintenance
* 18:30 root@cumin1003: START - Cookbook sre.mysql.pool pool db1251: Maintenance
* 18:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2022.codfw.wmnet with reason: host reimage
* 18:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1251 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95698 and previous config saved to /var/cache/conftool/dbconfig/20260729-182428-cwilliams.json
* 18:24 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2014.codfw.wmnet
* 18:24 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2014.codfw.wmnet
* 18:24 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1251.eqiad.wmnet with reason: Maintenance
* 18:23 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1235: Maintenance
* 18:22 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 18:19 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 18:19 brett@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 18:17 brett@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 18:17 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 18:16 mutante: removing jenkins during the train - living on the edge - no, just kidding, jenkins has migrated to dedicated machines, nothing should happen
* 18:15 brett@cumin2002: END (ERROR) - Cookbook sre.loadbalancer.restart-pybal (exit_code=97) rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]])
* 18:15 mutante: CI: contint1002/contint2002: apt-get remove --purge jenkins - jenkins be gone - [[phab:T418521|T418521]]
* 18:13 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on P<nowiki>{</nowiki>lvs2014.codfw.wmnet<nowiki>}</nowiki> and A:lvs ([[phab:T428495|T428495]])
* 18:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2022
* 18:08 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2022
* 18:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]]
* 18:03 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2022
* 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:02 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2022.codfw.wmnet 211.48.192.10.in-addr.arpa 1.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:02 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:02 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003"
* 18:02 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2022 - bking@cumin2003"
* 17:57 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:56 brett@cumin2002: END (FAIL) - Cookbook sre.loadbalancer.restart-pybal (exit_code=1) rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]])
* 17:55 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]]
* 17:54 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2022
* 17:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2022.codfw.wmnet with OS bookworm
* 17:50 brett@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-codfw and A:lvs ([[phab:T428495|T428495]])
* 17:47 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2015.codfw.wmnet -> wdqs2021.codfw.wmnet, repooling source-only afterwards
* 17:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s)
* 17:47 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients back to codfw - [[phab:T428495|T428495]]
* 17:47 swfrench@dns1004: END - running authdns-update
* 17:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1252 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95692 and previous config saved to /var/cache/conftool/dbconfig/20260729-174713-cwilliams.json
* 17:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1252.eqiad.wmnet with reason: Maintenance
* 17:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1249: Maintenance
* 17:45 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2015\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 17:45 swfrench@dns1004: START - running authdns-update
* 17:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db1202: Maintenance
* 17:41 akhatun: Deployed refinery using scap, then deployed onto hdfs
* 17:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1202 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95688 and previous config saved to /var/cache/conftool/dbconfig/20260729-173759-cwilliams.json
* 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1202.eqiad.wmnet with reason: Maintenance
* 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1194: Maintenance
* 17:37 root@cumin1003: START - Cookbook sre.mysql.pool pool db1235: Maintenance
* 17:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1230: Maintenance
* 17:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1235 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95684 and previous config saved to /var/cache/conftool/dbconfig/20260729-173051-cwilliams.json
* 17:30 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1235.eqiad.wmnet with reason: Maintenance
* 17:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1234: Maintenance
* 17:26 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674] (duration: 02m 02s)
* 17:24 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (thin): Regular analytics weekly train THIN [analytics/refinery@56695674]
* 17:23 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674] (duration: 06m 20s)
* 17:20 dancy@deploy1003: Finished scap sync-world: Testing delay_messageblobstore_purge: true (duration: 06m 29s)
* 17:17 akhatun@deploy1003: Started deploy [analytics/refinery@5669567]: Regular analytics weekly train [analytics/refinery@56695674]
* 17:17 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 00m 22s)
* 17:16 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674]
* 17:13 dancy@deploy1003: Started scap sync-world: Testing delay_messageblobstore_purge: true
* 17:05 mutante: CI: contint1002/contint2002 - restarted httpd to be extra sure all is cleaned up - https://integration.wikimedia.org/ci/ is up and running [[phab:T418521|T418521]]
* 17:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 17:03 mutante: CI: contint1002/contint2002 - rm /etc/apache2/jenkins_proxy - removing legacy jenkins proxy config - jenkins is on new dedicated machines and uses jenkins_proxy_ext config [[phab:T418521|T418521]]
* 17:02 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] (duration: 36m 25s)
* 17:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1249: Maintenance
* 16:59 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 16:54 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards
* 16:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1249 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95674 and previous config saved to /var/cache/conftool/dbconfig/20260729-165339-cwilliams.json
* 16:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1249.eqiad.wmnet with reason: Maintenance
* 16:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1248: Maintenance
* 16:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1194: Maintenance
* 16:47 swfrench-wmf: silenced EtcdReplicationDown 57b2b421-1cc9-4e38-9276-{{Gerrit|94f223fd231c}} - [[phab:T428495|T428495]]
* 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 16:46 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye
* 16:46 tchin@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/eventstreams-internal: apply
* 16:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1230: Maintenance
* 16:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95669 and previous config saved to /var/cache/conftool/dbconfig/20260729-164422-cwilliams.json
* 16:44 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 16:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1194.eqiad.wmnet with reason: Maintenance
* 16:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1191: Maintenance
* 16:43 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1003.eqiad.wmnet
* 16:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1234: Maintenance
* 16:43 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 16:43 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Rolling back deployment
* 16:43 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host an-test-master1004.eqiad.wmnet
* 16:41 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 16:40 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 16:40 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 16:39 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1230 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95667 and previous config saved to /var/cache/conftool/dbconfig/20260729-163932-cwilliams.json
* 16:39 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 16:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1230.eqiad.wmnet with reason: Maintenance
* 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1207: Maintenance
* 16:38 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 16:37 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host an-test-master1004.eqiad.wmnet
* 16:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1234 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95664 and previous config saved to /var/cache/conftool/dbconfig/20260729-163719-cwilliams.json
* 16:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1234.eqiad.wmnet with reason: Maintenance
* 16:36 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1079.eqiad.wmnet with OS trixie
* 16:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1232: Maintenance
* 16:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1259: Maintenance
* 16:28 akhatun@deploy1003: Finished deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674] (duration: 06m 57s)
* 16:28 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:26 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318707{{!}}Optimize language name loading with fallbacks (T231755)]]
* 16:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1051 hosts
* 16:21 akhatun@deploy1003: Started deploy [analytics/refinery@5669567] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@56695674]
* 16:20 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2006.codfw.wmnet with OS bookworm
* 16:19 akhatun: Deploying Refinery at {{Gerrit|56695674}} as part of weekly train
* 16:18 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage
* 16:16 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] (duration: 15m 36s)
* 16:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2021.codfw.wmnet with OS bookworm
* 16:14 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1079.eqiad.wmnet with reason: host reimage
* 16:12 topranks: hot-swap line card in FPC0 on cr1-eqiad with replacement MPC10E from Juniper [[phab:T426343|T426343]]
* 16:10 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment
* 16:07 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1248: Maintenance
* 16:01 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318705{{!}}Load language names from JSON instead of PHP (T231755)]], [[gerrit:1318706{{!}}Update rebuild.php to also write message JSON files (T231755)]]
* 16:00 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 16:00 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 15:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95651 and previous config saved to /var/cache/conftool/dbconfig/20260729-155956-cwilliams.json
* 15:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1248.eqiad.wmnet with reason: Maintenance
* 15:59 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage
* 15:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1247: Maintenance
* 15:59 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2074.codfw.wmnet with OS trixie
* 15:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1191: Maintenance
* 15:57 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:55 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1079.eqiad.wmnet with OS trixie
* 15:55 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2006.codfw.wmnet with reason: host reimage
* 15:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1207: Maintenance
* 15:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95646 and previous config saved to /var/cache/conftool/dbconfig/20260729-155104-cwilliams.json
* 15:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1191.eqiad.wmnet with reason: Maintenance
* 15:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1181: Maintenance
* 15:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db1232: Maintenance
* 15:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage
* 15:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1207 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95643 and previous config saved to /var/cache/conftool/dbconfig/20260729-154735-cwilliams.json
* 15:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1259: Maintenance
* 15:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1207.eqiad.wmnet with reason: Maintenance
* 15:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1200: Maintenance
* 15:46 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] (duration: 31m 59s)
* 15:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 15:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1232 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95640 and previous config saved to /var/cache/conftool/dbconfig/20260729-154330-cwilliams.json
* 15:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1232.eqiad.wmnet with reason: Maintenance
* 15:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1219: Maintenance
* 15:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2015.codfw.wmnet, repooling source-only afterwards
* 15:41 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 18s)
* 15:41 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1259 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95638 and previous config saved to /var/cache/conftool/dbconfig/20260729-154107-cwilliams.json
* 15:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1259.eqiad.wmnet with reason: Maintenance
* 15:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1254: Maintenance
* 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2015.codfw.wmnet with OS bookworm
* 15:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2021.codfw.wmnet with reason: host reimage
* 15:36 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2006.codfw.wmnet with OS bookworm
* 15:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 15:35 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Continuing with deployment
* 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie
* 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 15:32 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:29 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:28 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ms-be2074.codfw.wmnet with OS trixie
* 15:28 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2006.codfw.wmnet
* 15:26 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1078.eqiad.wmnet with OS trixie
* 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2098.codfw.wmnet with reason: host reimage
* 15:25 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2097.codfw.wmnet with reason: host reimage
* 15:22 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2006.codfw.wmnet
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2021
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2021
* 15:19 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2021
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:19 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2021.codfw.wmnet 210.48.192.10.in-addr.arpa 0.1.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003"
* 15:19 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2021 - bking@cumin2003"
* 15:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1318704{{!}}Add i18n/LanguageNames to MessagesDirs (T231755)]]
* 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage
* 15:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db1247: Maintenance
* 15:11 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2004.codfw.wmnet with reason: [[phab:T433478|T433478]]
* 15:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2015.codfw.wmnet with reason: host reimage
* 15:10 klausman@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 8:00:00 on ml-serve2002.codfw.wmnet with reason: [[phab:T433476|T433476]]
* 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 15:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1247 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95625 and previous config saved to /var/cache/conftool/dbconfig/20260729-150459-cwilliams.json
* 15:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1247.eqiad.wmnet with reason: Maintenance
* 15:04 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 15:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Maintenance
* 15:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage
* 15:03 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org
* 15:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1181: Maintenance
* 15:01 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2098.codfw.wmnet with OS bullseye
* 15:00 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host ms-be2097.codfw.wmnet with OS bullseye
* 15:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db1200: Maintenance
* 14:59 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2074.codfw.wmnet with reason: host reimage
* 14:59 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1285.eqiad.wmnet with OS trixie
* 14:59 Amir1: mwscript-k8s -- extensions/TimedMediaHandler/maintenance/requeueTranscodes.php --wiki=commonswiki --key '360p.mpeg4.mov' --throttle --video --missing ([[phab:T358266|T358266]])
* 14:58 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org
* 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2098']
* 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2098']
* 14:58 jhancock@cumin2002: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['ms-be2097']
* 14:58 jhancock@cumin2002: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['ms-be2097']
* 14:58 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1078.eqiad.wmnet with reason: host reimage
* 14:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95621 and previous config saved to /var/cache/conftool/dbconfig/20260729-145629-cwilliams.json
* 14:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1181.eqiad.wmnet with reason: Maintenance
* 14:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1174: Maintenance
* 14:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1219: Maintenance
* 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1006.wikimedia.org on all recursors
* 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1006.wikimedia.org on all recursors
* 14:55 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader1005.wikimedia.org on all recursors
* 14:55 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader1005.wikimedia.org on all recursors
* 14:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1200 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95618 and previous config saved to /var/cache/conftool/dbconfig/20260729-145336-cwilliams.json
* 14:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1254: Maintenance
* 14:53 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1200.eqiad.wmnet with reason: Maintenance
* 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1185: Maintenance
* 14:52 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2021
* 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2015
* 14:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2015
* 14:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95616 and previous config saved to /var/cache/conftool/dbconfig/20260729-144946-cwilliams.json
* 14:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1219.eqiad.wmnet with reason: Maintenance
* 14:49 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1218: Maintenance
* 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2021.codfw.wmnet with OS bookworm
* 14:48 dancy@deploy1003: Finished deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]]) (duration: 00m 15s)
* 14:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2015.codfw.wmnet with OS bookworm
* 14:48 dancy@deploy1003: Started deploy [zuul/deploy@22703a6]: Deploying https://gerrit.wikimedia.org/r/c/integration/zuul/+/1311501 ([[phab:T432491|T432491]])
* 14:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1254 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95613 and previous config saved to /var/cache/conftool/dbconfig/20260729-144729-cwilliams.json
* 14:47 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1254.eqiad.wmnet with reason: Maintenance
* 14:47 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1233: Maintenance
* 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2013\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 14:46 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2014\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 14:46 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage
* 14:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:42 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:40 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1285.eqiad.wmnet with reason: host reimage
* 14:39 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1078.eqiad.wmnet with OS trixie
* 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2074.codfw.wmnet with OS trixie
* 14:32 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:32 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:32 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:31 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2005.codfw.wmnet with OS bookworm
* 14:30 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:29 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:29 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:27 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1285
* 14:27 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1285
* 14:27 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1285.eqiad.wmnet with OS trixie
* 14:24 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 14:24 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:24 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:24 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:23 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:22 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:22 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:21 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:17 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 14:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Maintenance
* 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:15 pt1979@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002"
* 14:15 pt1979@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Add asw1-604 loopback ipv4 - pt1979@cumin2002"
* 14:12 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:12 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95599 and previous config saved to /var/cache/conftool/dbconfig/20260729-141014-cwilliams.json
* 14:10 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.update-replication (exit_code=1)
* 14:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1244.eqiad.wmnet with reason: Maintenance
* 14:10 pt1979@cumin2002: START - Cookbook sre.dns.netbox
* 14:10 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1243: Maintenance
* 14:09 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage
* 14:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1174: Maintenance
* 14:08 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad
* 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db1185: Maintenance
* 14:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 14:05 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2005.codfw.wmnet with reason: host reimage
* 14:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95595 and previous config saved to /var/cache/conftool/dbconfig/20260729-140309-cwilliams.json
* 14:03 sukhe@dns1004: END - running authdns-update
* 14:03 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1174.eqiad.wmnet with reason: Maintenance
* 14:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1170: Maintenance
* 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db1218: Maintenance
* 14:01 sukhe@dns1004: START - running authdns-update
* 14:00 sukhe@dns1004: START - running authdns-update
* 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db1233: Maintenance
* 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1185 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95592 and previous config saved to /var/cache/conftool/dbconfig/20260729-135925-cwilliams.json
* 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1185.eqiad.wmnet with reason: Maintenance
* 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1161: Maintenance
* 13:58 sukhe@puppetserver1001: conftool action : set/pooled=true; selector: dnsdisc=urldownloader
* 13:58 root@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1265.eqiad.wmnet with OS trixie
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95590 and previous config saved to /var/cache/conftool/dbconfig/20260729-135621-cwilliams.json
* 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1218.eqiad.wmnet with reason: Maintenance
* 13:55 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1206: Maintenance
* 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2073.codfw.wmnet with OS trixie
* 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1233 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95587 and previous config saved to /var/cache/conftool/dbconfig/20260729-135335-cwilliams.json
* 13:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1233.eqiad.wmnet with reason: Maintenance
* 13:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1229: Maintenance
* 13:50 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/kartotherian: apply
* 13:50 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service
* 13:49 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:49 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/kartotherian: apply
* 13:48 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 13:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1077.eqiad.wmnet with OS trixie
* 13:47 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 13:46 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2005.codfw.wmnet with OS bookworm
* 13:44 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 13:44 root@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage
* 13:40 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] (duration: 09m 22s)
* 13:39 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:38 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:36 root@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1265.eqiad.wmnet with reason: host reimage
* 13:35 stran@deploy1003: stran: Continuing with deployment
* 13:33 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 13:32 stran@deploy1003: stran: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified t
* 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage
* 13:30 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ml-build1001.eqiad.wmnet
* 13:30 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1319079{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319080{{!}}SI: Instrument interaction timeline link (T433053)]], [[gerrit:1319077{{!}}SI: Instrument abuse filter hits link (T433053)]], [[gerrit:1319078{{!}}SI: Instrument abuse filter hits link (T433053)]]
* 13:29 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad
* 13:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage
* 13:27 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 13:27 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:27 sukhe@cumin1003: END (FAIL) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=99) for role: url_downloader@eqiad
* 13:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2073.codfw.wmnet with reason: host reimage
* 13:26 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] (duration: 07m 56s)
* 13:25 klausman@cumin1003: START - Cookbook sre.hosts.reboot-single for host ml-build1001.eqiad.wmnet
* 13:24 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 13:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet
* 13:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet
* 13:24 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1077.eqiad.wmnet with reason: host reimage
* 13:23 root@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host db1265
* 13:23 root@cumin1003: START - Cookbook sre.hosts.move-vlan for host db1265
* 13:23 root@cumin1003: START - Cookbook sre.hosts.reimage for host db1265.eqiad.wmnet with OS trixie
* 13:23 root@cumin1003: START - Cookbook sre.mysql.pool pool db1243: Maintenance
* 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2005.codfw.wmnet
* 13:22 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 13:22 samtar@deploy1003: dreamrimmer, samtar: Continuing with deployment
* 13:20 samtar@deploy1003: dreamrimmer, samtar: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.decommission (exit_code=1) for hosts an-test-master[1001-1002].eqiad.wmnet
* 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 13:19 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 13:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet
* 13:18 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=97) for role: url_downloader@eqiad
* 13:18 sukhe@cumin1003: START - Cookbook sre.loadbalancer.migrate-service-ipip for role: url_downloader@eqiad
* 13:18 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1309675{{!}}Enable the abuse filter block action on Hindi Wikipedia (T431830)]]
* 13:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95574 and previous config saved to /var/cache/conftool/dbconfig/20260729-131638-cwilliams.json
* 13:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1243.eqiad.wmnet with reason: Maintenance
* 13:16 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2005.codfw.wmnet
* 13:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1242: Maintenance
* 13:14 samtar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] (duration: 07m 00s)
* 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db1170: Maintenance
* 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet
* 13:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet
* 13:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet
* 13:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1223: Maintenance
* 13:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1161: Maintenance
* 13:10 samtar@deploy1003: anzx, samtar: Continuing with deployment
* 13:09 samtar@deploy1003: anzx, samtar: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db1206: Maintenance
* 13:08 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 13:07 samtar@deploy1003: Started scap sync-world: Backport for [[gerrit:1319075{{!}}remove throttle exceptions for concluded events]]
* 13:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95566 and previous config saved to /var/cache/conftool/dbconfig/20260729-130730-cwilliams.json
* 13:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1170.eqiad.wmnet with reason: Maintenance
* 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:07 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003"
* 13:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1158: Maintenance
* 13:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet
* 13:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95564 and previous config saved to /var/cache/conftool/dbconfig/20260729-130616-cwilliams.json
* 13:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 13:06 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1077.eqiad.wmnet with OS trixie
* 13:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db1229: Maintenance
* 13:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1161.eqiad.wmnet with reason: Maintenance
* 13:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt IPs new switches - cmooney@cumin1003"
* 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2073.codfw.wmnet with OS trixie
* 13:05 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1159: Maintenance
* 13:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95562 and previous config saved to /var/cache/conftool/dbconfig/20260729-130258-cwilliams.json
* 13:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1206.eqiad.wmnet with reason: Maintenance
* 13:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1196: Maintenance
* 13:01 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 13:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:00 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 13:00 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 12:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1229 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95559 and previous config saved to /var/cache/conftool/dbconfig/20260729-125950-cwilliams.json
* 12:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1229.eqiad.wmnet with reason: Maintenance
* 12:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1222: Maintenance
* 12:57 sukhe: sudo cumin 'A:lvs and (A:eqiad or A:codfw)' 'disable-puppet "adding new service urldownloader"': [[phab:T429175|T429175]]
* 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet
* 12:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet
* 12:56 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet
* 12:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet
* 12:50 sukhe: sudo cumin 'O:url_downloader' 'run-puppet-agent --enable "merging CR 1313948"': [[phab:T429175|T429175]]
* 12:48 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-master[1001-1002].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet
* 12:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet
* 12:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet
* 12:45 sukhe: sudo cumin 'O:url_downloader' 'disable-puppet "merging CR 1313948"': [[phab:T429175|T429175]]
* 12:44 btullis@cumin1003: START - Cookbook sre.dns.netbox
* 12:40 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet
* 12:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet
* 12:38 ayounsi@dns1004: END - running authdns-update
* 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet
* 12:35 ayounsi@dns1004: START - running authdns-update
* 12:34 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-master[1001-1002].eqiad.wmnet
* 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts an-test-coord1001.eqiad.wmnet
* 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:33 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet
* 12:32 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve101[2-5].eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 12:29 root@cumin1003: START - Cookbook sre.mysql.pool pool db1242: Maintenance
* 12:25 root@cumin1003: START - Cookbook sre.mysql.pool pool db1223: Maintenance
* 12:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95544 and previous config saved to /var/cache/conftool/dbconfig/20260729-122254-cwilliams.json
* 12:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1242.eqiad.wmnet with reason: Maintenance
* 12:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1241: Maintenance
* 12:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1051 hosts
* 12:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db1158: Maintenance
* 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95540 and previous config saved to /var/cache/conftool/dbconfig/20260729-121937-cwilliams.json
* 12:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1223.eqiad.wmnet with reason: Maintenance
* 12:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1212: Maintenance
* 12:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db1159: Maintenance
* 12:17 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/services/proton: sync
* 12:15 elukey@deploy1003: helmfile [eqiad] START helmfile.d/services/proton: sync
* 12:15 root@cumin1003: START - Cookbook sre.mysql.pool pool db1196: Maintenance
* 12:14 Daimona: Creating new DB tables for the CampaignEvents extension in x1.testwiki, x1.test2wiki, x1.officewiki, and x1.wikishared # [[phab:T429339|T429339]]
* 12:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db1222: Maintenance
* 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95535 and previous config saved to /var/cache/conftool/dbconfig/20260729-121211-cwilliams.json
* 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1158.eqiad.wmnet with reason: Maintenance
* 12:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1159 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95534 and previous config saved to /var/cache/conftool/dbconfig/20260729-121146-cwilliams.json
* 12:11 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1159.eqiad.wmnet with reason: Maintenance
* 12:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95533 and previous config saved to /var/cache/conftool/dbconfig/20260729-120847-cwilliams.json
* 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 12:08 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1196.eqiad.wmnet with reason: Maintenance
* 12:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1195: Maintenance
* 12:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95530 and previous config saved to /var/cache/conftool/dbconfig/20260729-120424-cwilliams.json
* 12:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1222.eqiad.wmnet with reason: Maintenance
* 12:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1098 hosts
* 12:00 marostegui: Rename tables [[phab:T425074|T425074]]
* 12:00 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: an-test-coord1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 11:58 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1197: Maintenance
* 11:55 btullis@cumin1003: START - Cookbook sre.dns.netbox
* 11:52 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/services/proton: sync
* 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:51 elukey@deploy1003: helmfile [codfw] START helmfile.d/services/proton: sync
* 11:51 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:50 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts an-test-coord1001.eqiad.wmnet
* 11:50 elukey@deploy1003: helmfile [staging] DONE helmfile.d/services/proton: sync
* 11:49 elukey@deploy1003: helmfile [staging] START helmfile.d/services/proton: sync
* 11:35 root@cumin1003: START - Cookbook sre.mysql.pool pool db1241: Maintenance
* 11:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db1212: Maintenance
* 11:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:29 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1241 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95520 and previous config saved to /var/cache/conftool/dbconfig/20260729-112918-cwilliams.json
* 11:29 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1241.eqiad.wmnet with reason: Maintenance
* 11:29 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1238: Maintenance
* 11:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1212 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95517 and previous config saved to /var/cache/conftool/dbconfig/20260729-112727-cwilliams.json
* 11:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 11:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1212.eqiad.wmnet with reason: Maintenance
* 11:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1198: Maintenance
* 11:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:22 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 11:21 root@cumin1003: START - Cookbook sre.mysql.pool pool db1195: Maintenance
* 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95514 and previous config saved to /var/cache/conftool/dbconfig/20260729-111450-cwilliams.json
* 11:14 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1195.eqiad.wmnet with reason: Maintenance
* 11:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1186: Maintenance
* 11:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 11:05 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:54 marostegui: Dropping renamed tables [[phab:T425066|T425066]]
* 10:41 root@cumin1003: START - Cookbook sre.mysql.pool pool db1238: Maintenance
* 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1198: Maintenance
* 10:39 Amir1: ran https://phabricator.wikimedia.org/T432509#12149723 in production ([[phab:T432509|T432509]])
* 10:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db1197: Maintenance
* 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95501 and previous config saved to /var/cache/conftool/dbconfig/20260729-103532-cwilliams.json
* 10:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1238.eqiad.wmnet with reason: Maintenance
* 10:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1221: Maintenance
* 10:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1198 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95499 and previous config saved to /var/cache/conftool/dbconfig/20260729-103330-cwilliams.json
* 10:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1198.eqiad.wmnet with reason: Maintenance
* 10:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1175: Maintenance
* 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1197 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95496 and previous config saved to /var/cache/conftool/dbconfig/20260729-103217-cwilliams.json
* 10:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1197.eqiad.wmnet with reason: Maintenance
* 10:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1188: Maintenance
* 10:27 root@cumin1003: START - Cookbook sre.mysql.pool pool db1186: Maintenance
* 10:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95493 and previous config saved to /var/cache/conftool/dbconfig/20260729-102111-cwilliams.json
* 10:21 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1186.eqiad.wmnet with reason: Maintenance
* 10:14 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 10:14 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 09:53 XioNoX: reboot cr2-magru - [[phab:T431750|T431750]]
* 09:52 XioNoX: drain cr2-magru - [[phab:T431750|T431750]]
* 09:48 root@cumin1003: START - Cookbook sre.mysql.pool pool db1221: Maintenance
* 09:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host zookeeper-test1002.eqiad.wmnet
* 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1188: Maintenance
* 09:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db1175: Maintenance
* 09:44 btullis@dns1004: END - running authdns-update
* 09:42 btullis@dns1004: START - running authdns-update
* 09:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95483 and previous config saved to /var/cache/conftool/dbconfig/20260729-094200-cwilliams.json
* 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 7 hosts with reason: Maintenance
* 09:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1221.eqiad.wmnet with reason: Maintenance
* 09:41 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host zookeeper-test1002.eqiad.wmnet
* 09:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1199: Maintenance
* 09:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95481 and previous config saved to /var/cache/conftool/dbconfig/20260729-093917-cwilliams.json
* 09:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1188.eqiad.wmnet with reason: Maintenance
* 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1182: Maintenance
* 09:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95479 and previous config saved to /var/cache/conftool/dbconfig/20260729-093842-cwilliams.json
* 09:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1175.eqiad.wmnet with reason: Maintenance
* 09:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1166: Maintenance
* 09:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1169: Maintenance
* 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s8
* 09:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1033.eqiad.wmnet,service=s5
* 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s8
* 09:33 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1033.eqiad.wmnet,service=s5
* 09:21 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:21 XioNoX: reboot cr1-magru - [[phab:T431750|T431750]]
* 09:17 XioNoX: drain cr1-magru - [[phab:T431750|T431750]]
* 09:15 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm1001.wikimedia.org
* 09:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/spark-history: apply
* 09:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/spark-history: apply
* 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 09:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 09:11 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-magru,cr2-magru IPv6,cr2-magru.mgmt with reason: router upgrade
* 09:11 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm1001.wikimedia.org
* 09:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp1005.wikimedia.org
* 09:07 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp1005.wikimedia.org
* 09:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp2005.wikimedia.org
* 09:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp2005.wikimedia.org
* 09:00 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade
* 09:00 marostegui: Dropping renamed tables [[phab:T426341|T426341]]
* 08:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1199: Maintenance
* 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1182: Maintenance
* 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db1166: Maintenance
* 08:47 root@cumin1003: START - Cookbook sre.mysql.pool pool db1169: Maintenance
* 08:46 ayounsi@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on cr1-magru,cr1-magru IPv6,cr1-magru.mgmt with reason: router upgrade
* 08:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2072.codfw.wmnet with OS trixie
* 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1199 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95464 and previous config saved to /var/cache/conftool/dbconfig/20260729-084534-cwilliams.json
* 08:45 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1199.eqiad.wmnet with reason: Maintenance
* 08:45 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: router upgrade, [[phab:T431750|T431750]]]
* 08:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1190: Maintenance
* 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95462 and previous config saved to /var/cache/conftool/dbconfig/20260729-084436-cwilliams.json
* 08:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1182.eqiad.wmnet with reason: Maintenance
* 08:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1156: Maintenance
* 08:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95460 and previous config saved to /var/cache/conftool/dbconfig/20260729-084400-cwilliams.json
* 08:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1166.eqiad.wmnet with reason: Maintenance
* 08:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1157: Maintenance
* 08:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95458 and previous config saved to /var/cache/conftool/dbconfig/20260729-084147-cwilliams.json
* 08:41 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1169.eqiad.wmnet with reason: Maintenance
* 08:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1163: Maintenance
* 08:30 btullis@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 11 hosts with reason: Replacing the namenodes
* 08:23 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage
* 08:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1098 hosts
* 08:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2072.codfw.wmnet with reason: host reimage
* 07:58 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2072.codfw.wmnet with OS trixie
* 07:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1190: Maintenance
* 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1156: Maintenance
* 07:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1157: Maintenance
* 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool db1163: Maintenance
* 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95444 and previous config saved to /var/cache/conftool/dbconfig/20260729-074930-cwilliams.json
* 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1190.eqiad.wmnet with reason: Maintenance
* 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95443 and previous config saved to /var/cache/conftool/dbconfig/20260729-074914-cwilliams.json
* 07:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95442 and previous config saved to /var/cache/conftool/dbconfig/20260729-074906-cwilliams.json
* 07:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1157.eqiad.wmnet with reason: Maintenance
* 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on 6 hosts with reason: Maintenance
* 07:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1156.eqiad.wmnet with reason: Maintenance
* 07:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95441 and previous config saved to /var/cache/conftool/dbconfig/20260729-074652-cwilliams.json
* 07:46 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1163.eqiad.wmnet with reason: Maintenance
* 07:46 ayounsi@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2034.codfw.wmnet
* 07:42 ayounsi@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2034.codfw.wmnet
* 07:42 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2232.codfw.wmnet with OS trixie
* 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:34 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:33 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:31 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:19 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2232.codfw.wmnet with reason: host reimage
* 07:15 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2232.codfw.wmnet with reason: host reimage
* 06:58 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2232.codfw.wmnet with OS trixie
* 06:58 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2232].codfw.wmnet with reason: Reimage
* 06:26 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1164.eqiad.wmnet with OS trixie
* 06:05 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage
* 06:01 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1164.eqiad.wmnet with reason: host reimage
* 05:47 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1164.eqiad.wmnet with OS trixie
* 05:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1164.eqiad.wmnet with reason: Reimage
== 2026-07-28 ==
* 22:50 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1138.eqiad.wmnet
* 22:50 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1138.eqiad.wmnet
* 22:49 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1138.eqiad.wmnet
* 22:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards
* 22:08 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards
* 22:03 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1006.eqiad.wmnet with OS trixie
* 20:58 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] (duration: 08m 19s)
* 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2020.codfw.wmnet -> wdqs2014.codfw.wmnet, repooling source-only afterwards
* 20:55 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2013.codfw.wmnet, repooling source-only afterwards
* 20:54 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:54 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 14s)
* 20:54 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 20:53 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 30s)
* 20:53 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 20:52 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:51 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:50 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318746{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T19006 T234299 T360814 T385317 T432477 T433018)]], [[gerrit:1318747{{!}}Bump wikimedia/parsoid to 0.24.0-a17 (T433018)]]
* 20:49 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:43 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 20:34 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] (duration: 06m 54s)
* 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:34 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:33 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs-next: apply
* 20:32 lerickson@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs-next: apply
* 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:31 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:30 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:30 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:29 arlolra@deploy1003: arlolra: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:27 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318753{{!}}Support older config values in $wgJsonConfigModels (T433008)]]
* 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2014.codfw.wmnet with OS bookworm
* 20:21 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 20:21 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:20 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 20:19 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:19 swfrench-wmf: switched etcd-mirror replication from conf2005 to conf2004 - [[phab:T428495|T428495]]
* 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:17 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:15 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] (duration: 08m 26s)
* 20:12 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:11 arlolra@deploy1003: anzx, arlolra: Continuing with deployment
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:11 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 20:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2013.codfw.wmnet with OS bookworm
* 20:09 arlolra@deploy1003: anzx, arlolra: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1318718{{!}}bolwiki: set correct version of logo (T429951)]], [[gerrit:1318756{{!}}Remove icon beside reporting link on desktop view (T433303)]], [[gerrit:1318754{{!}}Remove icon beside reporting link on desktop view (T433303)]]
* 20:07 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2216: Maintenance
* 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage
* 19:57 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:54 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:53 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:52 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:52 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2014.codfw.wmnet with reason: host reimage
* 19:49 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:48 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage
* 19:42 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1006.eqiad.wmnet with OS trixie
* 19:41 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2013.codfw.wmnet with reason: host reimage
* 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:39 lerickson@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 19:39 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 19:38 lerickson@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 19:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2014
* 19:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2014
* 19:29 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2014.codfw.wmnet with OS bookworm
* 19:28 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1006.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 19:27 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1006
* 19:26 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2012\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 19:26 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1006
* 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:26 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 19:25 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1006] - vriley@cumin1003"
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2013
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2013
* 19:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2013
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2013.codfw.wmnet 84.0.192.10.in-addr.arpa 4.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:21 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003"
* 19:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2013 - bking@cumin2003"
* 19:21 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 19:20 root@cumin1003: START - Cookbook sre.mysql.pool pool db2216: Maintenance
* 19:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2216 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95435 and previous config saved to /var/cache/conftool/dbconfig/20260728-191343-cwilliams.json
* 19:13 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2216.codfw.wmnet with reason: Maintenance
* 19:13 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2203: Maintenance
* 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1005.eqiad.wmnet with OS trixie
* 19:06 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 19:06 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - vriley@cumin1003"
* 18:46 dduvall@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 18:45 bking@cumin2003: START - Cookbook sre.dns.netbox
* 18:45 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:43 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:40 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:36 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-eqiad: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 18:35 dancy@deploy1003: Installation of scap version "4.275.0" completed for 3 hosts
* 18:33 dancy@deploy1003: Installing scap version "4.275.0" for 3 host(s)
* 18:32 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2098.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:32 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:30 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org [reason: pool for all services after reimaging]
* 18:29 sukhe@dns1004: END - running authdns-update
* 18:27 sukhe@dns1004: START - running authdns-update
* 18:27 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns3003.wikimedia.org,service=authdns-update [reason: pool authdns-update after reimaging]
* 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2203: Maintenance
* 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95430 and previous config saved to /var/cache/conftool/dbconfig/20260728-181958-cwilliams.json
* 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2203.codfw.wmnet with reason: Maintenance
* 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: Maintenance
* 18:18 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host ms-be2097.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2098
* 18:17 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2098
* 18:17 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-be2097
* 18:16 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host ms-be2097
* 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:15 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002"
* 18:15 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding ms-be2097-8 to codfw - jhancock@cumin2002"
* 18:10 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 18:08 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 18:05 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns3003.wikimedia.org with OS trixie
* 18:03 vriley@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1005.eqiad.wmnet with reason: host reimage
* 17:56 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 17:45 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1005.eqiad.wmnet with OS trixie
* 17:45 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:41 sukhe@dns1004: END - running authdns-update
* 17:39 sukhe@dns1004: START - running authdns-update
* 17:36 sukhe@puppetserver1001: conftool action : set/weight=1; selector: cluster=urldownloader,service=squid
* 17:36 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1005.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 17:35 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: cluster=urldownloader,service=squid
* 17:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 17:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005
* 17:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 17:34 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005
* 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:34 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003"
* 17:34 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1005] - vriley@cumin1003"
* 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2188: Maintenance
* 17:29 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2188 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95425 and previous config saved to /var/cache/conftool/dbconfig/20260728-172609-cwilliams.json
* 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2188.codfw.wmnet with reason: Maintenance
* 17:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2176: Maintenance
* 17:19 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1138.eqiad.wmnet with OS trixie
* 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1005
* 17:18 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1005
* 17:18 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:15 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 17:13 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns3003.wikimedia.org with reason: host reimage
* 17:07 sukhe@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns3003.wikimedia.org with reason: host reimage
* 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-eqiad
* 17:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1015.eqiad.wmnet
* 17:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1015.eqiad.wmnet
* 16:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage
* 16:55 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:restbase-codfw: Upgrade Java to 17.0.20 — [[phab:T433028|T433028]] - eevans@cumin1003
* 16:54 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1138.eqiad.wmnet with reason: host reimage
* 16:53 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1015.eqiad.wmnet
* 16:43 sukhe@cumin1003: START - Cookbook sre.hosts.reimage for host dns3003.wikimedia.org with OS trixie
* 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1015.eqiad.wmnet
* 16:43 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1014.eqiad.wmnet
* 16:43 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1014.eqiad.wmnet
* 16:43 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=dns3003.wikimedia.org [reason: depooling for reimage to trixie]
* 16:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 290 hosts
* 16:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2176: Maintenance
* 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1138
* 16:38 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1138
* 16:37 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1138
* 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:37 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1138.eqiad.wmnet 193.32.64.10.in-addr.arpa 3.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:37 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003"
* 16:37 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1138 - jiji@cumin1003"
* 16:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1014.eqiad.wmnet
* 16:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2176 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95420 and previous config saved to /var/cache/conftool/dbconfig/20260728-163235-cwilliams.json
* 16:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2176.codfw.wmnet with reason: Maintenance
* 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1014.eqiad.wmnet
* 16:32 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1013.eqiad.wmnet
* 16:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1013.eqiad.wmnet
* 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2174: Maintenance
* 16:28 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards
* 16:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1013.eqiad.wmnet
* 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1013.eqiad.wmnet
* 16:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1012.eqiad.wmnet
* 16:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1012.eqiad.wmnet
* 16:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1012.eqiad.wmnet
* 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1012.eqiad.wmnet
* 16:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1011.eqiad.wmnet
* 16:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1011.eqiad.wmnet
* 16:00 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host conf2004.codfw.wmnet with OS bookworm
* 15:59 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1011.eqiad.wmnet
* 15:56 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 15:55 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1011.eqiad.wmnet
* 15:54 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1010.eqiad.wmnet
* 15:54 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1010.eqiad.wmnet
* 15:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 15:50 ladsgroup@deploy1003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 15:49 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1010.eqiad.wmnet
* 15:48 ladsgroup@deploy1003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 15:48 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 15:46 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2248: Maintenance
* 15:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2174: Maintenance
* 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1010.eqiad.wmnet
* 15:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1009.eqiad.wmnet
* 15:44 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1009.eqiad.wmnet
* 15:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1138
* 15:41 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1138.eqiad.wmnet with OS trixie
* 15:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1009.eqiad.wmnet
* 15:39 robh@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on arclamp2001.codfw.wmnet with reason: ram upgrade
* 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2174 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95413 and previous config saved to /var/cache/conftool/dbconfig/20260728-153844-cwilliams.json
* 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2174.codfw.wmnet with reason: Maintenance
* 15:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2173: Maintenance
* 15:37 ladsgroup@deploy1003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 15:35 ladsgroup@deploy1003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1009.eqiad.wmnet
* 15:34 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1008.eqiad.wmnet
* 15:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1008.eqiad.wmnet
* 15:31 ladsgroup@deploy1003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 15:31 ladsgroup@deploy1003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 15:29 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1008.eqiad.wmnet
* 15:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 290 hosts
* 15:25 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2013
* 15:25 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2195: Maintenance
* 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1008.eqiad.wmnet
* 15:24 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1007.eqiad.wmnet
* 15:24 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1007.eqiad.wmnet
* 15:22 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage
* 15:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2013.codfw.wmnet with OS bookworm
* 15:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 15:19 swfrench@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on conf2004.codfw.wmnet with reason: host reimage
* 15:19 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 15:17 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1007.eqiad.wmnet
* 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1007.eqiad.wmnet
* 15:12 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1006.eqiad.wmnet
* 15:12 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1006.eqiad.wmnet
* 15:11 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1138.eqiad.wmnet
* 15:11 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003
* 15:11 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1138.eqiad.wmnet
* 15:11 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1138.eqiad.wmnet
* 15:10 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]] (duration: 00m 43s)
* 15:10 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab1004 for [[phab:T433382|T433382]]
* 15:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts
* 15:09 brennen@deploy1003: Finished deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]] (duration: 00m 55s)
* 15:08 brennen@deploy1003: Started deploy [phabricator/deployment@f8b349f]: deploy phab2003 for [[phab:T433382|T433382]]
* 15:07 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2012.codfw.wmnet, repooling source-only afterwards
* 15:07 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts
* 15:06 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1137.eqiad.wmnet
* 15:06 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1137.eqiad.wmnet
* 15:06 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1137.eqiad.wmnet
* 15:05 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1006.eqiad.wmnet
* 15:05 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1022\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main
* 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab2003.codfw.wmnet with reason: deployment
* 15:01 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1005.eqiad.wmnet with reason: deployment
* 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1006.eqiad.wmnet
* 15:00 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1006.eqiad.wmnet with reason: deployment
* 15:00 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1005.eqiad.wmnet
* 15:00 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1005.eqiad.wmnet
* 14:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2248: Maintenance
* 14:59 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on phab1004.eqiad.wmnet with reason: deployment
* 14:59 swfrench@cumin2002: START - Cookbook sre.hosts.reimage for host conf2004.codfw.wmnet with OS bookworm
* 14:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1005.eqiad.wmnet
* 14:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2248 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95403 and previous config saved to /var/cache/conftool/dbconfig/20260728-145532-cwilliams.json
* 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on db[2245-2247].codfw.wmnet with reason: Maintenance
* 14:55 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2248.codfw.wmnet with reason: Maintenance
* 14:54 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Maintenance
* 14:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2173: Maintenance
* 14:51 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts
* 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1005.eqiad.wmnet
* 14:50 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1004.eqiad.wmnet
* 14:50 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1004.eqiad.wmnet
* 14:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts
* 14:45 swfrench@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host conf2004.codfw.wmnet
* 14:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2173 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95399 and previous config saved to /var/cache/conftool/dbconfig/20260728-144453-cwilliams.json
* 14:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2173.codfw.wmnet with reason: Maintenance
* 14:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2170: Maintenance
* 14:44 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1004.eqiad.wmnet
* 14:39 swfrench@cumin2002: START - Cookbook sre.hosts.reboot-single for host conf2004.codfw.wmnet
* 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1004.eqiad.wmnet
* 14:38 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 14:38 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 14:38 root@cumin1003: START - Cookbook sre.mysql.pool pool db2195: Maintenance
* 14:36 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards
* 14:36 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 14:33 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 14:33 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2222: Maintenance
* 14:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2195 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95394 and previous config saved to /var/cache/conftool/dbconfig/20260728-143218-cwilliams.json
* 14:32 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2195.codfw.wmnet with reason: Maintenance
* 14:31 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2181: Maintenance
* 14:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 14:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 14:25 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 14:25 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 14:25 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 14:23 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1002.eqiad.wmnet
* 14:23 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1002.eqiad.wmnet
* 14:23 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s)
* 14:23 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:18 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1002.eqiad.wmnet
* 14:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2012.codfw.wmnet with OS bookworm
* 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1002.eqiad.wmnet
* 14:13 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1001.eqiad.wmnet
* 14:13 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet
* 14:11 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 14:11 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 14:08 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet
* 14:07 elukey@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'sync'.
* 14:07 elukey@deploy1003: helmfile [eqiad] START helmfile.d/admin 'sync'.
* 14:06 elukey@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'sync'.
* 14:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Maintenance
* 14:06 XioNoX: un-drain cr2-esams - [[phab:T431751|T431751]]
* 14:05 elukey@deploy1003: helmfile [codfw] START helmfile.d/admin 'sync'.
* 14:02 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet
* 14:02 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-eqiad
* 14:01 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 14:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2240 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95384 and previous config saved to /var/cache/conftool/dbconfig/20260728-140011-cwilliams.json
* 14:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2240.codfw.wmnet with reason: Maintenance
* 13:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2237: Maintenance
* 13:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2170: Maintenance
* 13:56 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T428495|T428495]])
* 13:55 XioNoX: reboot cr2-esams - [[phab:T431751|T431751]]
* 13:52 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 13:51 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr2-esams,cr2-esams IPv6,cr2-esams.mgmt with reason: router upgrade
* 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T428495|T428495]])
* 13:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2170 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95381 and previous config saved to /var/cache/conftool/dbconfig/20260728-135043-cwilliams.json
* 13:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2170.codfw.wmnet with reason: Maintenance
* 13:50 XioNoX: drain cr2-esams - [[phab:T431751|T431751]]
* 13:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2153: Maintenance
* 13:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2222: Maintenance
* 13:45 sukhe: restart pybal on A:lvs-codfw
* 13:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2012.codfw.wmnet with reason: host reimage
* 13:45 root@cumin1003: START - Cookbook sre.mysql.pool pool db2181: Maintenance
* 13:44 btullis@dns1004: END - running authdns-update
* 13:42 sukhe: restart pybal on lvs2014
* 13:42 btullis@dns1004: START - running authdns-update
* 13:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2222 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95376 and previous config saved to /var/cache/conftool/dbconfig/20260728-133948-cwilliams.json
* 13:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2222.codfw.wmnet with reason: Maintenance
* 13:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2221: Maintenance
* 13:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2181 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95374 and previous config saved to /var/cache/conftool/dbconfig/20260728-133857-cwilliams.json
* 13:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2181.codfw.wmnet with reason: Maintenance
* 13:38 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2167: Maintenance
* 13:30 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T428495|T428495]]
* 13:29 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 13:29 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431751|T431751]]]
* 13:29 ayounsi@cumin1003: END (FAIL) - Cookbook sre.dns.admin (exit_code=99) DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]]
* 13:28 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool esams [reason: router upgrade, [[phab:T431749|T431749]]]
* 13:27 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1022.eqiad.wmnet, repooling source-only afterwards
* 13:27 swfrench-wmf: begin rolling restart of confd in codfw, eqsin, ulsfo - [[phab:T428495|T428495]]
* 13:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2012
* 13:27 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2012
* 13:21 lucaswerkmeister-wmde@deploy1003: mwscript-k8s job started: cleanupTitles bolwiki # [[phab:T429951|T429951]]
* 13:21 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] (duration: 07m 19s)
* 13:20 swfrench-wmf: authdns-update to direct codfw, eqsin, ulsfo etcd clients to eqiad - [[phab:T428495|T428495]]
* 13:18 swfrench@dns1004: END - running authdns-update
* 13:17 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Continuing with deployment
* 13:16 swfrench@dns1004: START - running authdns-update
* 13:16 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2012
* 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:16 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2012.codfw.wmnet 57.48.192.10.in-addr.arpa 7.5.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:16 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:16 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, anzx: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[2025-2038].codfw.wmnet,restbase[1032-1045].eqiad.wmnet: Upgrade Cassandra to 5.0.8 — [[phab:T433028|T433028]] - eevans@cumin1003
* 13:14 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315908{{!}}magwiki: update logos (T433011)]], [[gerrit:1313960{{!}}bolwiki: add logo, sitename, projectnamespace and timezone (T429951)]]
* 13:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance
* 13:13 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2237: Maintenance
* 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003"
* 13:12 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback IPV6 for asw1-604 - pt1979@cumin2003"
* 13:11 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 20s)
* 13:11 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 13:10 esanders@deploy1003: Finished scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] (duration: 08m 11s)
* 13:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox
* 13:07 root@cumin1003: START - Cookbook sre.mysql.pool pool db2237: Maintenance
* 13:06 esanders@deploy1003: esanders: Continuing with deployment
* 13:04 esanders@deploy1003: esanders: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance
* 13:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2153: Maintenance
* 13:02 esanders@deploy1003: Started scap sync-world: Backport for [[gerrit:1309727{{!}}Set wgMFFallbackEditor to 'visual' on enwiki (T431858)]]
* 13:01 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95362 and previous config saved to /var/cache/conftool/dbconfig/20260728-130107-cwilliams.json
* 13:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2237.codfw.wmnet with reason: Maintenance
* 13:00 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2236: Maintenance
* 12:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2153: Maintenance
* 12:52 root@cumin1003: START - Cookbook sre.mysql.pool pool db2221: Maintenance
* 12:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2153 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95358 and previous config saved to /var/cache/conftool/dbconfig/20260728-125214-cwilliams.json
* 12:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2153.codfw.wmnet with reason: Maintenance
* 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2167: Maintenance
* 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-worker-codfw
* 12:51 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2011.codfw.wmnet
* 12:51 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2011.codfw.wmnet
* 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:49 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003"
* 12:48 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add loopback for asw1-603 - pt1979@cumin2003"
* 12:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2011.codfw.wmnet
* 12:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2221 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95357 and previous config saved to /var/cache/conftool/dbconfig/20260728-124601-cwilliams.json
* 12:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2221.codfw.wmnet with reason: Maintenance
* 12:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2218: Maintenance
* 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2167 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95354 and previous config saved to /var/cache/conftool/dbconfig/20260728-124457-cwilliams.json
* 12:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2167.codfw.wmnet with reason: Maintenance
* 12:44 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2166: Maintenance
* 12:42 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/kartotherian: apply
* 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2011.codfw.wmnet
* 12:41 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2010.codfw.wmnet
* 12:41 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2010.codfw.wmnet
* 12:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2010.codfw.wmnet
* 12:34 pt1979@cumin2003: START - Cookbook sre.dns.netbox
* 12:32 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 12:32 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 12:32 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2010.codfw.wmnet
* 12:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet
* 12:31 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet
* 12:27 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet
* 12:22 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet
* 12:22 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2008.codfw.wmnet
* 12:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2008.codfw.wmnet
* 12:16 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-604-eqsin
* 12:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2008.codfw.wmnet
* 12:16 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-604-eqsin
* 12:14 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance
* 12:14 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2236: Maintenance
* 12:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1137.eqiad.wmnet with OS trixie
* 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2008.codfw.wmnet
* 12:11 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2007.codfw.wmnet
* 12:11 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2007.codfw.wmnet
* 12:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2236: Maintenance
* 12:06 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2007.codfw.wmnet
* 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2236 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95348 and previous config saved to /var/cache/conftool/dbconfig/20260728-120253-cwilliams.json
* 12:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2236.codfw.wmnet with reason: Maintenance
* 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2007.codfw.wmnet
* 12:01 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet
* 12:01 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet
* 11:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2218: Maintenance
* 11:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db2166: Maintenance
* 11:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2219: Maintenance
* 11:56 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet
* 11:52 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage
* 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2218 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95344 and previous config saved to /var/cache/conftool/dbconfig/20260728-115155-cwilliams.json
* 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2218.codfw.wmnet with reason: Maintenance
* 11:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2208: Maintenance
* 11:51 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2166 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95342 and previous config saved to /var/cache/conftool/dbconfig/20260728-115119-cwilliams.json
* 11:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2166.codfw.wmnet with reason: Maintenance
* 11:50 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: Maintenance
* 11:47 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1137.eqiad.wmnet with reason: host reimage
* 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet
* 11:45 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2005.codfw.wmnet
* 11:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2005.codfw.wmnet
* 11:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2005.codfw.wmnet
* 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2005.codfw.wmnet
* 11:35 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2004.codfw.wmnet
* 11:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2004.codfw.wmnet
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1137
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1137
* 11:30 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1137
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:30 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1137.eqiad.wmnet 192.32.64.10.in-addr.arpa 2.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:30 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003"
* 11:25 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2004.codfw.wmnet
* 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2004.codfw.wmnet
* 11:19 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2003.codfw.wmnet
* 11:19 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2003.codfw.wmnet
* 11:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2003.codfw.wmnet
* 11:11 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance
* 11:10 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2219: Maintenance
* 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2219: Maintenance
* 11:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2003.codfw.wmnet
* 11:04 root@cumin1003: START - Cookbook sre.mysql.pool pool db2164: Maintenance
* 11:03 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet
* 11:03 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet
* 11:03 root@cumin1003: START - Cookbook sre.mysql.pool pool db2208: Maintenance
* 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2164 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95332 and previous config saved to /var/cache/conftool/dbconfig/20260728-105749-cwilliams.json
* 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2164.codfw.wmnet with reason: Maintenance
* 10:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2208 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95331 and previous config saved to /var/cache/conftool/dbconfig/20260728-105711-cwilliams.json
* 10:57 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2208.codfw.wmnet with reason: Maintenance
* 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2219 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95330 and previous config saved to /var/cache/conftool/dbconfig/20260728-105652-cwilliams.json
* 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2219.codfw.wmnet with reason: Maintenance
* 10:53 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1137 - jiji@cumin1003"
* 10:52 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet
* 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet
* 10:47 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2001.codfw.wmnet
* 10:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2001.codfw.wmnet
* 10:39 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2001.codfw.wmnet
* 10:35 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 10:34 jforrester@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] (duration: 09m 31s)
* 10:34 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1137
* 10:34 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2001.codfw.wmnet
* 10:34 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-worker-codfw
* 10:34 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1137.eqiad.wmnet with OS trixie
* 10:28 jforrester@deploy1003: jforrester: Continuing with deployment
* 10:27 jforrester@deploy1003: jforrester: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:25 jforrester@deploy1003: Started scap sync-world: Backport for [[gerrit:1307506{{!}}logging: Switch the wmfconfig processor to Monolog 3's type (LogRecord) (T397070)]]
* 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 10:21 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 10:21 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1137.eqiad.wmnet
* 10:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 10:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1137.eqiad.wmnet
* 10:20 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1137.eqiad.wmnet
* 10:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: Maintenance
* 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-staging-worker
* 09:37 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet
* 09:37 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet
* 09:32 kevinbazira@deploy1003: helmfile [ml-serve-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:31 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet
* 09:30 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 09:30 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance
* 09:30 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 09:30 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:30 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:22 klausman@cumin1003: END (ERROR) - Cookbook sre.ganeti.reboot-vm (exit_code=97) for VM ml-serve-ctrl2001.codfw.wmnet
* 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet
* 09:22 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-eqiad
* 09:22 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-eqiad
* 09:21 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet
* 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2002.codfw.wmnet
* 09:20 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2002.codfw.wmnet
* 09:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f2-codfw
* 09:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f2-codfw
* 09:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e4-codfw
* 09:18 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2163: Maintenance
* 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e4-codfw
* 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e2-codfw
* 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e2-codfw
* 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-e5-codfw
* 09:17 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-e5-codfw
* 09:17 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-f4-codfw
* 09:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2210: Maintenance
* 09:16 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-f4-codfw
* 09:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2182: Maintenance
* 09:14 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2002.codfw.wmnet
* 09:12 root@cumin1003: START - Cookbook sre.mysql.pool pool db2163: Maintenance
* 09:11 XioNoX: rebooting cr2-drmrs - [[phab:T431749|T431749]]
* 09:10 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr2-drmrs,cr2-drmrs IPv6,cr2-drmrs.mgmt with reason: router upgrade
* 09:06 XioNoX: draining cr2-drmrs - [[phab:T431749|T431749]]
* 09:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2163 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95320 and previous config saved to /var/cache/conftool/dbconfig/20260728-090638-cwilliams.json
* 09:06 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2163.codfw.wmnet with reason: Maintenance
* 09:06 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: Maintenance
* 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2002.codfw.wmnet
* 09:04 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2001.codfw.wmnet
* 09:04 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2001.codfw.wmnet
* 08:57 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2001.codfw.wmnet
* 08:48 XioNoX: un-drain cr1-drmrs - [[phab:T431749|T431749]]
* 08:47 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2001.codfw.wmnet
* 08:47 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-staging-worker
* 08:42 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:35 XioNoX: rebooting cr1-drmrs - [[phab:T431749|T431749]]
* 08:33 XioNoX: draining cr1-drmrs - [[phab:T431749|T431749]]
* 08:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance
* 08:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance
* 08:21 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2182: Maintenance
* 08:17 root@cumin1003: START - Cookbook sre.mysql.pool pool db2161: Maintenance
* 08:16 root@cumin1003: START - Cookbook sre.mysql.pool pool db2182: Maintenance
* 08:12 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2210: Maintenance
* 08:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2161 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95309 and previous config saved to /var/cache/conftool/dbconfig/20260728-081044-cwilliams.json
* 08:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2161.codfw.wmnet with reason: Maintenance
* 08:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2154: Maintenance
* 08:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2182 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95307 and previous config saved to /var/cache/conftool/dbconfig/20260728-080947-cwilliams.json
* 08:09 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2182.codfw.wmnet with reason: Maintenance
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2168: Maintenance
* 08:06 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-drmrs,cr1-drmrs IPv6,cr1-drmrs.mgmt with reason: router upgrade
* 08:06 root@cumin1003: START - Cookbook sre.mysql.pool pool db2210: Maintenance
* 08:05 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 08:05 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool drmrs [reason: router upgrade, [[phab:T431749|T431749]]]
* 08:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2210 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95305 and previous config saved to /var/cache/conftool/dbconfig/20260728-080008-cwilliams.json
* 08:00 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2210.codfw.wmnet with reason: Maintenance
* 07:59 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2206: Maintenance
* 07:50 gkyziridis@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 07:50 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2154: Maintenance
* 07:22 root@cumin1003: START - Cookbook sre.mysql.pool pool db2168: Maintenance
* 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2154 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95295 and previous config saved to /var/cache/conftool/dbconfig/20260728-071640-cwilliams.json
* 07:16 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2154.codfw.wmnet with reason: Maintenance
* 07:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95294 and previous config saved to /var/cache/conftool/dbconfig/20260728-071604-cwilliams.json
* 07:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2168.codfw.wmnet with reason: Maintenance
* 07:08 root@cumin1003: START - Cookbook sre.mysql.pool pool db2206: Maintenance
* 07:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2206 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95292 and previous config saved to /var/cache/conftool/dbconfig/20260728-070219-cwilliams.json
* 07:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2206.codfw.wmnet with reason: Maintenance
* 06:44 marostegui: Failover m5 from db1164 to db1228 - [[phab:T432967|T432967]]
* 06:39 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2235].codfw.wmnet,db[1164,1217,1228].eqiad.wmnet with reason: m5 master switch [[phab:T432967|T432967]]
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.10 (duration: 02m 34s)
* 03:39 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]] (duration: 36m 06s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.13 refs [[phab:T430832|T430832]]
* 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host zuul1004.eqiad.wmnet with OS trixie
* 02:57 dzahn@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003"
* 02:55 dzahn@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - dzahn@cumin1003"
* 02:37 dzahn@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage
* 02:31 dzahn@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on zuul1004.eqiad.wmnet with reason: host reimage
* 02:16 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie
* 02:15 dzahn@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host zuul1004.eqiad.wmnet with OS trixie
* 01:43 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie
* 01:43 dzahn@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS trixie
* 01:25 pt1979@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-603-eqsin
* 01:24 pt1979@cumin1003: START - Cookbook sre.network.tls for network device asw1-603-eqsin
* 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 01:12 pt1979@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003"
* 01:12 pt1979@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add mgmt for new switches in eqsin - pt1979@cumin2003"
* 01:08 pt1979@cumin2003: START - Cookbook sre.dns.netbox
* 00:48 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 00:47 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 00:47 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 00:47 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 00:26 mutante: attempting reimage with trixie on zuul1004 re-purposed physical hardware - dcops reported install issue - host was in busybox shell ([[phab:T427353|T427353]])
* 00:24 dzahn@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS trixie
== 2026-07-27 ==
* 23:50 Amir1: mass deleting vp8 transcodes
* 23:28 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:27 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host zuul1004.eqiad.wmnet with OS bullseye
* 23:26 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/blunderbuss: apply
* 23:25 amastilovic@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/blunderbuss: apply
* 22:39 maryum: Deploy security fix for [[phab:T432877|T432877]]
* 22:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1022.eqiad.wmnet with OS bookworm
* 22:37 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bullseye
* 22:32 sbassett: Deployed security fix for [[phab:T432789|T432789]]
* 22:22 sbassett: Deployed security patch for [[phab:T431819|T431819]]
* 22:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage
* 22:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1022.eqiad.wmnet with reason: host reimage
* 22:01 RScout-WMF: Deployed security fix for [[phab:T431819|T431819]]
* 22:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2012
* 21:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2012.codfw.wmnet with OS bookworm
* 21:55 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2011\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 21:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1022
* 21:45 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1022
* 21:44 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1022
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:44 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1022.eqiad.wmnet 239.48.64.10.in-addr.arpa 9.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:41 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:41 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 21:34 vriley@cumin1003: START - Cookbook sre.hosts.reimage for host zuul1004.eqiad.wmnet with OS bookworm
* 21:31 vriley@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:22 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:21 vriley@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:19 vriley@cumin1003: START - Cookbook sre.hosts.provision for host zuul1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART and with Dell SCP reboot policy FORCED
* 21:17 vriley@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host zuul1004
* 21:16 vriley@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host zuul1004
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:15 vriley@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
* 21:15 vriley@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update mgmt [zuul1004] - vriley@cumin1003"
* 21:11 vriley@cumin1003: START - Cookbook sre.dns.netbox
* 21:10 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2011.codfw.wmnet, repooling source-only afterwards
* 21:05 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:01 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1022
* 20:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1022.eqiad.wmnet with OS bookworm
* 20:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1021.eqiad.wmnet, repooling source-only afterwards
* 20:51 mutante: zuul1001 - re-enabled puppet - revert "cherry-picked" gerrit:1314120 - [[phab:T431003|T431003]]
* 20:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Maintenance
* 20:15 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] (duration: 08m 03s)
* 20:11 sbisson@deploy1003: sbisson: Continuing with deployment
* 20:09 sbisson@deploy1003: sbisson: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1318149{{!}}Article Guidance: clean up old wikidata config (T421250)]], [[gerrit:1318217{{!}}ArticleGuidance: Enable wgArticleGuidanceWikidataConnectEnabled in prod (T421250)]]
* 19:47 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s)
* 19:47 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Maintenance
* 19:27 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] (duration: 12m 26s)
* 19:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2228 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95285 and previous config saved to /var/cache/conftool/dbconfig/20260727-192711-cwilliams.json
* 19:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2228.codfw.wmnet with reason: Maintenance
* 19:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2223: Maintenance
* 19:23 krinkle@deploy1003: krinkle: Continuing with deployment
* 19:16 krinkle@deploy1003: krinkle: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:15 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1316171{{!}}logging: Simplify $wmgEnableExtraLogFile without $wmgExtraLogFile]], [[gerrit:1316172{{!}}logging: Remove $wmgUdp2logDest duplicate in favor of $wmgLocalServices]]
* 19:12 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2238: Maintenance
* 18:58 swfrench@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 18:57 swfrench@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 18:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2227: Maintenance
* 18:57 swfrench@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-experimental: apply
* 18:55 swfrench@deploy1003: helmfile [codfw] START helmfile.d/services/mw-experimental: apply
* 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1021.eqiad.wmnet with OS bookworm
* 18:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2011.codfw.wmnet with OS bookworm
* 18:40 root@cumin1003: START - Cookbook sre.mysql.pool pool db2223: Maintenance
* 18:39 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] (duration: 07m 05s)
* 18:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2223 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95275 and previous config saved to /var/cache/conftool/dbconfig/20260727-183500-cwilliams.json
* 18:34 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2223.codfw.wmnet with reason: Maintenance
* 18:34 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 18:34 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2213: Maintenance
* 18:33 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:32 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1318211{{!}}CodeMirrorWikiEditor: don't autofocus from live preview when RTP is open]]
* 18:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db2238: Maintenance
* 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage
* 18:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2238 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95271 and previous config saved to /var/cache/conftool/dbconfig/20260727-181944-cwilliams.json
* 18:19 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2238.codfw.wmnet with reason: Maintenance
* 18:19 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2226: Maintenance
* 18:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage
* 18:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2011.codfw.wmnet with reason: host reimage
* 18:12 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1021.eqiad.wmnet with reason: host reimage
* 18:09 root@cumin1003: START - Cookbook sre.mysql.pool pool db2227: Maintenance
* 18:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2227 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95265 and previous config saved to /var/cache/conftool/dbconfig/20260727-180256-cwilliams.json
* 18:02 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2227.codfw.wmnet with reason: Maintenance
* 18:02 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2194: Maintenance
* 17:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2011
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2011
* 17:56 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2011
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:56 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2011.codfw.wmnet 37.32.192.10.in-addr.arpa 7.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003"
* 17:56 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2011 - bking@cumin2003"
* 17:52 bking@cumin2003: START - Cookbook sre.dns.netbox
* 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2011
* 17:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1021
* 17:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1021
* 17:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2011.codfw.wmnet with OS bookworm
* 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1021.eqiad.wmnet with OS bookworm
* 17:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Maintenance
* 17:38 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2010\.codfw\.wmnet,dc=codfw,cluster=wdqs\-main,service=wdqs\-main
* 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95260 and previous config saved to /var/cache/conftool/dbconfig/20260727-173740-cwilliams.json
* 17:37 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance
* 17:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2211: Maintenance
* 17:36 bking@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1020\.eqiad\.wmnet,dc=eqiad,cluster=wdqs\-main,service=wdqs\-main
* 17:32 root@cumin1003: START - Cookbook sre.mysql.pool pool db2226: Maintenance
* 17:31 taavi@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] (duration: 06m 33s)
* 17:27 taavi@deploy1003: taavi: Continuing with deployment
* 17:27 taavi@deploy1003: taavi: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:26 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2226 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95256 and previous config saved to /var/cache/conftool/dbconfig/20260727-172636-cwilliams.json
* 17:26 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2226.codfw.wmnet with reason: Maintenance
* 17:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2225: Maintenance
* 17:25 taavi@deploy1003: Started scap sync-world: Backport for [[gerrit:1318202{{!}}Undeploy WP25EasterEggs (I) (T418134)]]
* 17:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 17:11 root@cumin1003: START - Cookbook sre.mysql.pool pool db2194: Maintenance
* 17:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2194 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95248 and previous config saved to /var/cache/conftool/dbconfig/20260727-170453-cwilliams.json
* 17:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2194.codfw.wmnet with reason: Maintenance
* 17:04 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2190: Maintenance
* 16:52 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 16:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2211: Maintenance
* 16:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95242 and previous config saved to /var/cache/conftool/dbconfig/20260727-164015-cwilliams.json
* 16:40 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2211.codfw.wmnet with reason: Maintenance
* 16:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: Maintenance
* 16:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db2225: Maintenance
* 16:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2172: Maintenance
* 16:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 16:38 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 16:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2225 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95238 and previous config saved to /var/cache/conftool/dbconfig/20260727-163307-cwilliams.json
* 16:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2225.codfw.wmnet with reason: Maintenance
* 16:32 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: Maintenance
* 16:13 root@cumin1003: START - Cookbook sre.mysql.pool pool db2190: Maintenance
* 16:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2190 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95230 and previous config saved to /var/cache/conftool/dbconfig/20260727-160602-cwilliams.json
* 16:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2190.codfw.wmnet with reason: Maintenance
* 15:53 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2177: Maintenance
* 15:53 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance
* 15:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db2178: Maintenance
* 15:51 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2172: Maintenance
* 15:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2178 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95224 and previous config saved to /var/cache/conftool/dbconfig/20260727-154559-cwilliams.json
* 15:46 root@cumin1003: START - Cookbook sre.mysql.pool pool db2172: Maintenance
* 15:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2178.codfw.wmnet with reason: Maintenance
* 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2171: Maintenance
* 15:44 root@cumin1003: START - Cookbook sre.mysql.pool pool db2189: Maintenance
* 15:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1026.eqiad.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:41 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2172 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95222 and previous config saved to /var/cache/conftool/dbconfig/20260727-153927-cwilliams.json
* 15:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2172.codfw.wmnet with reason: Maintenance
* 15:38 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:38 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2189 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95220 and previous config saved to /var/cache/conftool/dbconfig/20260727-153833-cwilliams.json
* 15:38 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2189.codfw.wmnet with reason: Maintenance
* 15:34 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:32 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 15:32 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 15:31 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:29 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:26 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2010.codfw.wmnet, repooling source-only afterwards
* 15:22 zabe@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] (duration: 07m 00s)
* 15:21 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1020.eqiad.wmnet, repooling source-only afterwards
* 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2155: Maintenance
* 15:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2175: Maintenance
* 15:18 zabe@deploy1003: zabe: Continuing with deployment
* 15:17 zabe@deploy1003: zabe: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:15 zabe@deploy1003: Started scap sync-world: Backport for [[gerrit:1318166{{!}}SpecialWantedFiles: Simplify query plan (T431518)]]
* 15:15 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 06s)
* 15:15 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2010.codfw.wmnet with OS bookworm
* 15:08 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance
* 15:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2177: Maintenance
* 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2177: Maintenance
* 14:58 root@cumin1003: START - Cookbook sre.mysql.pool pool db2171: Maintenance
* 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2171 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95209 and previous config saved to /var/cache/conftool/dbconfig/20260727-145236-cwilliams.json
* 14:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2171.codfw.wmnet with reason: Maintenance
* 14:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2177 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95208 and previous config saved to /var/cache/conftool/dbconfig/20260727-145206-cwilliams.json
* 14:52 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2157: Maintenance
* 14:51 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2177.codfw.wmnet with reason: Maintenance
* 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1020.eqiad.wmnet with OS bookworm
* 14:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2156: Maintenance
* 14:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage
* 14:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2010.codfw.wmnet with reason: host reimage
* 14:41 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003
* 14:34 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance
* 14:33 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2010
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2010
* 14:24 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2010
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2010.codfw.wmnet 94.16.192.10.in-addr.arpa 4.9.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:24 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003"
* 14:24 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2010 - bking@cumin2003"
* 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage
* 14:23 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching restbase[1031,2024]*: Upgrade Cassandra to 5.0.8 (canary) - eevans@cumin1003
* 14:20 jgiannelos@deploy1003: helmfile [staging] DONE helmfile.d/services/kartotherian: apply
* 14:20 jgiannelos@deploy1003: helmfile [staging] START helmfile.d/services/kartotherian: apply
* 14:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1020.eqiad.wmnet with reason: host reimage
* 14:17 sukhe: sudo gnt-instance reboot urldownloader1005.wikimedia.org
* 14:16 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:15 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:14 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:13 jelto@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:08 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2155: Maintenance
* 14:05 root@cumin1003: START - Cookbook sre.mysql.pool pool db2157: Maintenance
* 14:04 root@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2175: Maintenance
* 14:03 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 11 hosts
* 14:02 root@cumin1003: START - Cookbook sre.mysql.pool pool db2155: Maintenance
* 14:01 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 11 hosts
* 14:01 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1136.eqiad.wmnet
* 14:01 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1136.eqiad.wmnet
* 14:01 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1136.eqiad.wmnet
* 14:00 root@cumin1003: START - Cookbook sre.mysql.pool pool db2156: Maintenance
* 13:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2157 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95194 and previous config saved to /var/cache/conftool/dbconfig/20260727-135943-cwilliams.json
* 13:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2157.codfw.wmnet with reason: Maintenance
* 13:59 root@cumin1003: START - Cookbook sre.mysql.pool pool db2175: Maintenance
* 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:58 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 13:57 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2010
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2155 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95193 and previous config saved to /var/cache/conftool/dbconfig/20260727-135613-cwilliams.json
* 13:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2155.codfw.wmnet with reason: Maintenance
* 13:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1020
* 13:55 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1020
* 13:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2156 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95192 and previous config saved to /var/cache/conftool/dbconfig/20260727-135413-cwilliams.json
* 13:54 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2156.codfw.wmnet with reason: Maintenance
* 13:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2175 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95191 and previous config saved to /var/cache/conftool/dbconfig/20260727-135300-cwilliams.json
* 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2010.codfw.wmnet with OS bookworm
* 13:52 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2175.codfw.wmnet with reason: Maintenance
* 13:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1020.eqiad.wmnet with OS bookworm
* 13:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 34 hosts
* 13:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 34 hosts
* 13:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1201: Maintenance
* 13:27 Lucas_WMDE: UTC afternoon backport+config window doen
* 13:18 lucaswerkmeister-wmde@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] (duration: 11m 57s)
* 13:14 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Continuing with deployment
* 13:08 lucaswerkmeister-wmde@deploy1003: lucaswerkmeister-wmde, sihe: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:07 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]]
* 13:07 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool ulsfo [reason: router upgrade finished, [[phab:T431752|T431752]]]
* 13:06 lucaswerkmeister-wmde@deploy1003: Started scap sync-world: Backport for [[gerrit:1315843{{!}}Enable WikibaseLexeme REST API on beta wikidata (T430943)]], [[gerrit:1315842{{!}}Remove obsolete wmgWikibaseRestApiEnabled setting (T302959 T324999 T383774)]]
* 13:03 XioNoX: repool cr4-ulsfo - [[phab:T431752|T431752]]
* 12:51 root@cumin1003: START - Cookbook sre.mysql.pool pool db1201: Maintenance
* 12:48 gkyziridis@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'edit-check' for release 'main' .
* 12:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1201 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95186 and previous config saved to /var/cache/conftool/dbconfig/20260727-124404-cwilliams.json
* 12:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1201.eqiad.wmnet with reason: Maintenance
* 12:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1187: Maintenance
* 12:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) url-downloader.eqiad.wikimedia.org on all recursors
* 12:30 sukhe@cumin1003: START - Cookbook sre.dns.wipe-cache url-downloader.eqiad.wikimedia.org on all recursors
* 12:30 sukhe@dns1004: END - running authdns-update
* 12:30 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] (duration: 09m 32s)
* 12:28 sukhe@dns1004: START - running authdns-update
* 12:25 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 12:22 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:20 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1318098{{!}}Enable CheckUser SI special page on dewiki and ukwiki (T432835 T433226)]]
* 12:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts
* 12:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts
* 12:13 XioNoX: rebooting cr4-ulsfo for upgrade - [[phab:T431752|T431752]]
* 12:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: es1038 repool
* 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 38 hosts
* 12:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 38 hosts
* 11:55 root@cumin1003: START - Cookbook sre.mysql.pool pool db1187: Maintenance
* 11:53 urbanecm@deploy1003: mwscript-k8s job started: foreachwikiindblist growthexperiments GrowthExperiments:cleanMentorList # [[phab:T431804|T431804]]
* 11:50 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr4-ulsfo,cr4-ulsfo IPv6,cr4-ulsfo.mgmt with reason: router upgrade
* 11:50 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] (duration: 11m 07s)
* 11:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1187 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95178 and previous config saved to /var/cache/conftool/dbconfig/20260727-114844-cwilliams.json
* 11:48 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1187.eqiad.wmnet with reason: Maintenance
* 11:43 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 11:42 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:39 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1311832{{!}}[Growth] Deploy automated mentor list cleaner to all wikis (T431804)]]
* 11:37 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:36 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 11:36 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:35 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 11:29 XioNoX: start draining cr4-ulsfo - [[phab:T431752|T431752]]
* 11:29 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 11:29 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 11:28 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1035: testing
* 11:28 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1035: testing
* 11:27 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1035: testing
* 11:27 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1035: testing
* 11:26 ayounsi@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]]
* 11:26 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: es1038 repool
* 11:26 ayounsi@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool ulsfo [reason: router upgrade, [[phab:T431752|T431752]]]
* 11:26 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing
* 11:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1264: Maintenance
* 11:24 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing
* 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050 as master', diff saved to https://phabricator.wikimedia.org/P95170 and previous config saved to /var/cache/conftool/dbconfig/20260727-112326-marostegui.json
* 11:23 marostegui@cumin1003: dbctl commit (dc=all): 'Repool es1050', diff saved to https://phabricator.wikimedia.org/P95169 and previous config saved to /var/cache/conftool/dbconfig/20260727-112302-marostegui.json
* 11:22 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1050: testing
* 11:22 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1050: testing
* 11:20 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 11:18 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 11:18 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 11:12 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 11:11 blake@deploy1003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 11:09 blake@deploy1003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 11:09 btullis@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 11:08 blake@deploy1003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 11:05 blake@deploy1003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 11:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 11:02 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:50 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 10:43 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:39 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1264: Maintenance
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 10:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 10:37 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 10:36 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 10:35 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 10:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 10:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 10:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1264 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95164 and previous config saved to /var/cache/conftool/dbconfig/20260727-103204-cwilliams.json
* 10:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1264.eqiad.wmnet with reason: Maintenance
* 10:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 10:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 10:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 10:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 10:24 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1237: Maintenance
* 10:04 elukey: restart burrow main-eqiad on kafkamon2003 to clear some errors on kafka-main1008
* 09:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1136.eqiad.wmnet with OS trixie
* 09:39 elukey: restart burrow-main-eqiad.service on kafkamon1003 to see if a recurrent kafka error on kafka-main1008 goes away
* 09:39 root@cumin1003: START - Cookbook sre.mysql.pool pool db1237: Maintenance
* 09:38 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage
* 09:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1237 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95159 and previous config saved to /var/cache/conftool/dbconfig/20260727-093328-cwilliams.json
* 09:33 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1237.eqiad.wmnet with reason: Maintenance
* 09:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1136.eqiad.wmnet with reason: host reimage
* 09:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1203: Maintenance
* 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1136
* 09:17 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1136
* 09:04 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1136
* 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:04 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1136.eqiad.wmnet 191.32.64.10.in-addr.arpa 1.9.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:04 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003"
* 09:04 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1136 - jiji@cumin1003"
* 08:52 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 08:52 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 08:52 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 08:51 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 08:50 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 08:47 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1136
* 08:46 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1136.eqiad.wmnet with OS trixie
* 08:44 marostegui: Rename tables on s3 [[phab:T425066|T425066]]
* 08:43 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1136.eqiad.wmnet
* 08:43 root@cumin1003: START - Cookbook sre.mysql.pool pool db1203: Maintenance
* 08:43 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1136.eqiad.wmnet
* 08:43 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1136.eqiad.wmnet
* 08:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1203 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95154 and previous config saved to /var/cache/conftool/dbconfig/20260727-083703-cwilliams.json
* 08:36 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1203.eqiad.wmnet with reason: Maintenance
* 08:16 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1179: Maintenance
* 07:44 phuedx: UTC morning backport window done
* 07:37 phuedx@deploy1003: Finished scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] (duration: 32m 33s)
* 07:28 root@cumin1003: START - Cookbook sre.mysql.pool pool db1179: Maintenance
* 07:26 marostegui: Rename tables on s3 [[phab:T426341|T426341]]
* 07:25 phuedx@deploy1003: phuedx: Continuing with deployment
* 07:22 marostegui: Drop tables in akwiki nawiki pihwiki - growthexperiments_* [[phab:T428885|T428885]]
* 07:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1179 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95149 and previous config saved to /var/cache/conftool/dbconfig/20260727-072234-cwilliams.json
* 07:22 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1179.eqiad.wmnet with reason: Maintenance
* 07:20 phuedx@deploy1003: phuedx: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:16 ryankemper: [[phab:T430880|T430880]] [WDQS] Reimaged `wdqs1018` and `wdqs1019` to Bookworm, restored data using test-cookbook change {{Gerrit|1317128}}, and repooled both; 25/36 hosts complete
* 07:04 phuedx@deploy1003: Started scap sync-world: Backport for [[gerrit:1317814{{!}}sessionTick: Use Test Kitchen to send action=feature_not_available events (T413296)]]
* 06:57 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1019.eqiad.wmnet
* 06:56 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1018.eqiad.wmnet
* 06:40 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1020.eqiad.wmnet with reason: Cloning
* 06:35 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db1228.eqiad.wmnet with reason: Rebooting
* 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:29 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:25 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:25 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s6
* 06:25 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1032.eqiad.wmnet,service=s4
* 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards
* 06:00 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards
* 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1019.eqiad.wmnet, repooling source-only afterwards
* 04:51 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs1018.eqiad.wmnet, repooling source-only afterwards
* 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s)
* 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 04:48 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s)
* 04:48 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 36s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-26 ==
* 14:59 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:59 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:59 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:59 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:09 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1019.eqiad.wmnet with OS bookworm
* 01:05 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1018.eqiad.wmnet with OS bookworm
* 00:43 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage
* 00:38 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage
* 00:34 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1019.eqiad.wmnet with reason: host reimage
* 00:33 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1018.eqiad.wmnet with reason: host reimage
* 00:16 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 00:16 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 00:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 00:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1019
* 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1019
* 00:11 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1018
* 00:11 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1018
* 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1019.eqiad.wmnet with OS bookworm
* 00:08 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1018.eqiad.wmnet with OS bookworm
== 2026-07-25 ==
* 22:06 ryankemper: [[phab:T430880|T430880]] [WDQS] Repooled `wdqs1017` and `wdqs2024` after reimaging to bookworm, scap deploying, and data xfering
* 22:04 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2024.codfw.wmnet
* 22:03 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1017.eqiad.wmnet
* 21:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards
* 21:06 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards
* 20:52 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:52 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:52 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:52 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1017.eqiad.wmnet, repooling source-only afterwards
* 20:18 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2024.codfw.wmnet, repooling source-only afterwards
* 20:15 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:15 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:15 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:15 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 05s)
* 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 19:57 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s)
* 19:57 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 19:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2024.codfw.wmnet with OS bookworm
* 19:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1017.eqiad.wmnet with OS bookworm
* 19:02 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage
* 18:58 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage
* 18:53 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2024.codfw.wmnet with reason: host reimage
* 18:52 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1017.eqiad.wmnet with reason: host reimage
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2024
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2024
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1017
* 18:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1017
* 18:27 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2024
* 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:26 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2024.codfw.wmnet 58.16.192.10.in-addr.arpa 8.5.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:26 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:24 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1017
* 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:24 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1017.eqiad.wmnet 238.48.64.10.in-addr.arpa 8.3.2.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:24 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003"
* 18:24 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1017 - ryankemper@cumin2003"
* 18:23 ryankemper@cumin2003: START - Cookbook sre.dns.netbox
* 18:18 ryankemper@cumin2003: START - Cookbook sre.dns.netbox
* 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1017
* 18:17 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2024
* 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1017.eqiad.wmnet with OS bookworm
* 18:14 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2024.codfw.wmnet with OS bookworm
* 18:05 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs1016` and `wdqs2023` to Bookworm with `--move-vlan`, restored main and scholarly data, validated postflights, and repooled both hosts. Confirmed PyBal rebuilt both backends with their new addresses
* 17:45 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2023.codfw.wmnet
* 17:43 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1016.eqiad.wmnet
* 06:35 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards
* 06:17 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards
* 05:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs2023.codfw.wmnet, repooling source-only afterwards
* 05:19 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs1016.eqiad.wmnet, repooling source-only afterwards
* 05:07 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 07s)
* 05:07 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 05:06 ryankemper@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host (duration: 00m 06s)
* 05:06 ryankemper@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] deploy to fresh wdqs host
* 03:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2023.codfw.wmnet with OS bookworm
* 02:59 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage
* 02:56 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2023.codfw.wmnet with reason: host reimage
* 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2023
* 02:33 ryankemper@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2023
* 02:30 ryankemper@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2023
* 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 02:30 ryankemper@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2023.codfw.wmnet 35.0.192.10.in-addr.arpa 5.3.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 02:30 ryankemper@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003"
* 02:30 ryankemper@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2023 - ryankemper@cumin2003"
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:15 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1016.eqiad.wmnet with OS bookworm
* 00:49 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage
* 00:43 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1016.eqiad.wmnet with reason: host reimage
* 00:31 ryankemper@cumin2003: START - Cookbook sre.dns.netbox
* 00:27 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1016
* 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1016
* 00:27 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2023
* 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1016.eqiad.wmnet with OS bookworm
* 00:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2023.codfw.wmnet with OS bookworm
* 00:11 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs1014.eqiad.wmnet and wdqs2008.codfw.wmnet after Bookworm reimage, transfer, and postflight; wdqs2008 is serving, while wdqs1014 will remain outside of service until a pybal restart next monday
== 2026-07-24 ==
* 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1014.eqiad.wmnet
* 23:54 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2008.codfw.wmnet
* 23:43 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest2010.codfw.wmnet with OS trixie
* 23:08 jhathaway@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage
* 23:03 jhathaway@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest2010.codfw.wmnet with reason: host reimage
* 22:33 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 22:13 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:13 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:13 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:13 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:00 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 21:53 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 21:53 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie
* 21:51 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 21:47 jhathaway@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie
* 21:43 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 21:39 jhathaway@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 21:38 jhathaway@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 17:21 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1135.eqiad.wmnet
* 17:21 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1135.eqiad.wmnet
* 17:21 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1135.eqiad.wmnet
* 16:34 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 16:34 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 16:34 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 16:34 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 16:33 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 16:33 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 16:28 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:28 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:28 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:28 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:11 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards
* 15:56 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1135.eqiad.wmnet with OS trixie
* 15:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 40 hosts
* 15:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 40 hosts
* 15:37 topranks: upgrade SR-Linux OS on lswtest-d8-eqiad
* 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage
* 15:33 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: upgrade lswtest-d8-eqiad
* 15:32 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1135.eqiad.wmnet with reason: host reimage
* 15:30 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host mc-gp2006.codfw.wmnet with OS bookworm
* 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1135
* 15:15 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1135
* 15:13 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage
* 15:08 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on mc-gp2006.codfw.wmnet with reason: host reimage
* 14:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:48 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:34 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] (duration: 41m 12s)
* 14:32 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1135
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:32 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1135.eqiad.wmnet 177.32.64.10.in-addr.arpa 7.7.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003"
* 14:32 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1135 - jiji@cumin1003"
* 14:29 krinkle@deploy1003: krinkle: Continuing with deployment
* 14:29 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:27 jiji@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host mc-gp2006.codfw.wmnet with OS bookworm
* 14:26 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 14:15 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1135
* 14:14 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1135.eqiad.wmnet with OS trixie
* 14:14 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1135.eqiad.wmnet
* 14:13 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1135.eqiad.wmnet
* 14:13 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1135.eqiad.wmnet
* 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1072.eqiad.wmnet
* 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1072.eqiad.wmnet
* 14:10 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet
* 14:10 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet
* 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1071.eqiad.wmnet
* 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1071.eqiad.wmnet
* 14:09 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet
* 14:09 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet
* 13:58 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 13:58 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:58 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:57 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:55 krinkle@deploy1003: krinkle: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:53 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1315130{{!}}Add missing normaliseParams() call to ThreeDHandler::doTransform (T432911)]]
* 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:45 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003"
* 13:45 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push new IPs for mc-gp2006 - cmooney@cumin1003"
* 13:44 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) mc-gp2006.codfw.wmnet on all recursors
* 13:44 cmooney@cumin1003: START - Cookbook sre.dns.wipe-cache mc-gp2006.codfw.wmnet on all recursors
* 13:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host mc-gp2006.codfw.wmnet with OS bookworm
* 13:41 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:30 papaul: reboot mr1-eqsin for maintenance
* 13:24 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb[1029-1031].eqiad.wmnet
* 13:10 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb[1029-1031].eqiad.wmnet
* 11:33 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts
* 11:11 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts
* 10:56 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 7 hosts
* 10:47 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 7 hosts
* 10:44 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:44 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:41 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:41 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:35 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 8 hosts
* 10:34 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:33 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:32 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:32 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:31 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host sretest2010.codfw.wmnet with OS trixie
* 10:31 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:30 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 8 hosts
* 10:24 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host sretest2010.codfw.wmnet with OS trixie
* 10:19 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host sretest2010.codfw.wmnet with OS trixie
* 10:18 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 16 hosts
* 10:17 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2001.codfw.wmnet
* 10:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2001.codfw.wmnet
* 10:12 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2001.codfw.wmnet
* 10:02 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2001.codfw.wmnet
* 10:02 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-ctrl2002.codfw.wmnet
* 09:57 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-ctrl2002.codfw.wmnet
* 09:56 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl2002.codfw.wmnet
* 09:51 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl2002.codfw.wmnet
* 09:51 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1002.eqiad.wmnet
* 09:47 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1002.eqiad.wmnet
* 09:47 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-serve-ctrl1001.eqiad.wmnet
* 09:44 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-serve-ctrl1001.eqiad.wmnet
* 09:34 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2003.codfw.wmnet
* 09:32 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2003.codfw.wmnet
* 09:32 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2002.codfw.wmnet
* 09:30 brouberol@dns1004: END - running authdns-update
* 09:29 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2002.codfw.wmnet
* 09:29 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-staging-etcd2001.codfw.wmnet
* 09:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 9 hosts
* 09:27 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-staging-etcd2001.codfw.wmnet
* 09:27 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2001.codfw.wmnet
* 09:26 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 9 hosts
* 09:26 brouberol@dns1004: START - running authdns-update
* 09:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 57 hosts
* 09:24 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2001.codfw.wmnet
* 09:24 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2002.codfw.wmnet
* 09:22 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2002.codfw.wmnet
* 09:21 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 57 hosts
* 09:20 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd2003.codfw.wmnet
* 09:19 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 16 hosts
* 09:16 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd2003.codfw.wmnet
* 09:16 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1003.eqiad.wmnet
* 09:15 urbanecm@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 09:15 urbanecm@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 09:13 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1003.eqiad.wmnet
* 09:13 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1002.eqiad.wmnet
* 09:11 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1002.eqiad.wmnet
* 09:11 klausman@cumin1003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM ml-etcd1001.eqiad.wmnet
* 09:07 klausman@cumin1003: START - Cookbook sre.ganeti.reboot-vm for VM ml-etcd1001.eqiad.wmnet
* 08:32 trueg@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:24 trueg@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:16 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:07 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:07 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:07 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:02 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:01 trueg@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:59 trueg@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:57 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:57 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:32 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 06:46 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1025.eqiad.wmnet with reason: Cloning
* 06:46 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s4
* 06:45 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1025.eqiad.wmnet,service=s6
* 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s6
* 06:44 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1019.eqiad.wmnet,service=s4
* 03:40 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:40 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:40 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:40 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 03:37 ssastry@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:37 ssastry@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:37 ssastry@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:36 ssastry@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:49 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on mr1-eqsin,mr1-eqsin IPv6 with reason: connection issue
* 02:38 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on cr[2-3]-eqsin.mgmt,ps1-[603-604]-eqsin with reason: connection issue
* 02:08 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 27s)
* 02:01 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-23 ==
* 23:27 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on mr1-eqsin.oob,mr1-eqsin.oob IPv6 with reason: switch refresh
* 22:21 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003
* 22:01 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to NONE - eevans@cumin1003
* 21:29 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1014.eqiad.wmnet, repooling source-only afterwards
* 21:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 46s)
* 21:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 21:19 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003
* 21:00 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Setting storage compatibility to UPGRADING - eevans@cumin1003
* 20:17 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 57s)
* 20:13 dani@deploy1003: dani: Continuing with deployment
* 20:07 dani@deploy1003: dani: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:05 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314951{{!}}Undeploy Referring Experiences survey on enwiki (T432289)]]
* 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 19:24 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002"
* 19:24 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating the rest of the ipv6 dns records. - jhancock@cumin2002"
* 19:14 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 19:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs1014.eqiad.wmnet with OS bookworm
* 19:04 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.noop (exit_code=99)
* 19:04 cwilliams@cumin1003: START - Cookbook sre.mysql.noop
* 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage
* 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1014.eqiad.wmnet with reason: host reimage
* 18:30 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2007.codfw.wmnet -> wdqs2008.codfw.wmnet, repooling source-only afterwards
* 18:28 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 19s)
* 18:28 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1014
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1014
* 18:22 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1014
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:22 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1014.eqiad.wmnet 188.32.64.10.in-addr.arpa 8.8.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:22 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003"
* 18:21 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1014 - bking@cumin2003"
* 18:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2215: Maintenance
* 18:18 dancy@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 18:15 bking@cumin2003: START - Cookbook sre.dns.netbox
* 18:15 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:06 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 18:05 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 18:04 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2052: codfw rack B8 re-pool after maintenance
* 17:54 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:54 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 17:32 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2215: Maintenance
* 17:29 cmooney@dns3003: END - running authdns-update
* 17:27 cmooney@dns3003: START - running authdns-update
* 17:23 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new entries for row a b spines - cmooney@cumin1003"
* 17:22 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 17:21 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 17:18 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool es2052: codfw rack B8 re-pool after maintenance
* 17:18 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2189: codfw rack B8 re-pool after maintenance
* 17:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2215.codfw.wmnet with reason: Maintenance
* 17:17 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2215 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95126 and previous config saved to /var/cache/conftool/dbconfig/20260723-170903-cwilliams.json
* 17:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2191 to x1 primary [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95125 and previous config saved to /var/cache/conftool/dbconfig/20260723-170612-cwilliams.json
* 17:05 cezmunsta: Starting x1 codfw failover from db2215 to db2191 - [[phab:T432986|T432986]]
* 16:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2191 with weight 0 [[phab:T432986|T432986]]', diff saved to https://phabricator.wikimedia.org/P95123 and previous config saved to /var/cache/conftool/dbconfig/20260723-165831-cwilliams.json
* 16:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 16 hosts with reason: Primary switchover x1 [[phab:T432986|T432986]]
* 16:36 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 138128
* 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 138128
* 16:33 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2189: codfw rack B8 re-pool after maintenance
* 16:33 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2164: codfw rack B8 re-pool after maintenance
* 16:28 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1072.eqiad.wmnet
* 16:27 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1072.eqiad.wmnet with OS trixie
* 16:18 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2249: Maintenance
* 16:06 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage
* 16:06 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1072.eqiad.wmnet with reason: host reimage
* 15:50 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1072
* 15:50 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1072
* 15:49 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1072.eqiad.wmnet with OS trixie
* 15:48 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2164: codfw rack B8 re-pool after maintenance
* 15:48 musikanimal@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] (duration: 06m 37s)
* 15:48 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2163: codfw rack B8 re-pool after maintenance
* 15:45 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 15:43 musikanimal@deploy1003: musikanimal: Continuing with deployment
* 15:43 musikanimal@deploy1003: musikanimal: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:41 musikanimal@deploy1003: Started scap sync-world: Backport for [[gerrit:1314861{{!}}Util: Cast numeric config strings to their declared types (T432910)]]
* 15:36 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1072.eqiad.wmnet
* 15:35 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1072.eqiad.wmnet
* 15:35 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1072.eqiad.wmnet
* 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:34 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003"
* 15:34 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: push any outstanding updates - cmooney@cumin1003"
* 15:33 root@cumin1003: START - Cookbook sre.mysql.pool pool db2249: Maintenance
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 15:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:26 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 15:21 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:21 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 15:21 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:21 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 15:20 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 15:19 cmooney@dns2004: END - running authdns-update
* 15:17 cmooney@dns2004: START - running authdns-update
* 15:14 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns2004.wikimedia.org
* 15:12 brouberol@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 15:12 brouberol@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 15:12 klausman@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host ml-serve1001.eqiad.wmnet with OS trixie
* 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1071.eqiad.wmnet
* 15:11 jiji@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wikikube-worker1071.eqiad.wmnet with OS trixie
* 15:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=1) for host wdqs2008.codfw.wmnet with OS bookworm
* 15:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2249.codfw.wmnet with reason: Maintenance
* 15:08 brouberol@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 15:08 brouberol@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 15:08 brouberol@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 15:08 brouberol@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 15:06 cmooney@cumin1003: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org
* 15:02 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host kubestage2002.codfw.wmnet
* 15:02 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host kubestage2002.codfw.wmnet
* 15:02 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2163: codfw rack B8 re-pool after maintenance
* 15:01 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 15:01 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 14:59 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test2001.codfw.wmnet
* 14:57 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test2001.codfw.wmnet
* 14:56 ryankemper: [WDQS] [[phab:T430880|T430880]] Reimaged `wdqs2016` to Bookworm, xferred scholarly_articles from `wdqs2024`, validated updater/readiness/federation, and repooled
* 14:51 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2016.codfw.wmnet
* 14:51 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage
* 14:48 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage
* 14:47 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1001.eqiad.wmnet with reason: host reimage
* 14:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage
* 14:43 topranks: reboot lsw1-b8-codw to upgrade JunOS [[phab:T430929|T430929]]
* 14:41 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1071.eqiad.wmnet with reason: host reimage
* 14:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2008.codfw.wmnet with reason: host reimage
* 14:37 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2231: Maintenance
* 14:30 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1001.eqiad.wmnet with OS trixie
* 14:25 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host kubestage2002.codfw.wmnet
* 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1071
* 14:23 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1071
* 14:23 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 14:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards
* 14:22 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2052: codfw rack B8 depool for maintenance
* 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool es2052: codfw rack B8 depool for maintenance
* 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2249: codfw rack B8 depool for maintenance
* 14:21 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1071
* 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:21 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1071.eqiad.wmnet 166.48.64.10.in-addr.arpa 6.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:21 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003"
* 14:21 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1071 - jiji@cumin1003"
* 14:21 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2249: codfw rack B8 depool for maintenance
* 14:21 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2189: codfw rack B8 depool for maintenance
* 14:20 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host kubestage2002.codfw.wmnet
* 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2050.codfw.wmnet
* 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2189: codfw rack B8 depool for maintenance
* 14:20 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2164: codfw rack B8 depool for maintenance
* 14:20 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2164: codfw rack B8 depool for maintenance
* 14:19 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2163: codfw rack B8 depool for maintenance
* 14:19 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:19 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1014
* 14:19 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2163: codfw rack B8 depool for maintenance
* 14:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1014.eqiad.wmnet with OS bookworm
* 14:17 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2050.codfw.wmnet
* 14:16 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2035-2036,2087-2090,2286-2291].codfw.wmnet
* 14:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2008
* 14:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2008
* 14:14 cmooney@cumin1003: conftool action : set/pooled=no; selector: name=dns2004.wikimedia.org
* 14:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2008.codfw.wmnet with OS bookworm
* 14:13 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 14:12 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:10 topranks: depool dns2004 before lsw1-b8-codfw switch maintenance [[phab:T430929|T430929]]
* 14:10 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b8-codfw,lsw1-b8-codfw IPv6,lsw1-b8-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b8-codfw JunOS upgrade
* 14:07 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 30 hosts with reason: lsw1-b8-codfw JunOS upgrade
* 14:06 elukey: upload python3-docker-report 0.0.19 to apt.wikimedia.org for bookworm and trixie
* 13:59 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 13:58 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1071
* 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:57 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:57 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:57 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1071.eqiad.wmnet with OS trixie
* 13:55 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1071.eqiad.wmnet
* 13:55 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1071.eqiad.wmnet
* 13:55 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1071.eqiad.wmnet
* 13:53 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:52 logmsgbot: kharlan Deployed security patch for [[phab:T432948|T432948]]
* 13:51 tchin@deploy1003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 13:51 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:51 brouberol@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 13:51 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2231: Maintenance
* 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:50 tchin@deploy1003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 13:50 tchin@deploy1003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:49 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:49 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new dns entries for eqiad row a b new switches - cmooney@cumin1003"
* 13:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2231 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95097 and previous config saved to /var/cache/conftool/dbconfig/20260723-134436-cwilliams.json
* 13:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2231.codfw.wmnet with reason: Maintenance
* 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:39 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:38 brouberol@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 13:38 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] (duration: 09m 07s)
* 13:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1037 hosts
* 13:35 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2196: Maintenance
* 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:34 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:33 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 13:33 kharlan@deploy1003: kharlan, emc-wmf: Continuing with deployment
* 13:31 kharlan@deploy1003: kharlan, emc-wmf: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:30 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:28 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1311851{{!}}EventStreamConfig: remove stream used in past experiments (T428265)]]
* 13:17 hashar@deploy1003: Finished deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies (duration: 00m 14s)
* 13:17 hashar@deploy1003: Started deploy [integration/docroot@2199146]: build: License GPL2.0+ / updating npm dependencies
* 13:14 sukhe: sukhe@lvs1019:~$ sudo systemctl restart pybal.service
* 13:07 sukhe: sukhe@lvs1020:~$ sudo systemctl restart pybal.service
* 12:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2207: Repooling
* 12:49 root@cumin1003: START - Cookbook sre.mysql.pool pool db2196: Maintenance
* 12:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2196 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95087 and previous config saved to /var/cache/conftool/dbconfig/20260723-123952-cwilliams.json
* 12:39 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2196.codfw.wmnet with reason: Maintenance
* 12:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2191: Maintenance
* 12:29 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:13 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling
* 12:12 marostegui@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2207: Repooling
* 12:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool db2207: Repooling
* 11:56 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2235.codfw.wmnet with OS trixie
* 11:50 root@cumin1003: START - Cookbook sre.mysql.pool pool db2191: Maintenance
* 11:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1070.eqiad.wmnet
* 11:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1070.eqiad.wmnet
* 11:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1070.eqiad.wmnet
* 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2191 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95080 and previous config saved to /var/cache/conftool/dbconfig/20260723-114308-cwilliams.json
* 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2191.codfw.wmnet with reason: Maintenance
* 11:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2186: Maintenance
* 11:35 cmooney@cumin1003: END (PASS) - Cookbook sre.network.peering (exit_code=0) with action 'configure' for AS: 46375
* 11:34 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 46375
* 11:33 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2235.codfw.wmnet with reason: host reimage
* 11:28 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2235.codfw.wmnet with reason: host reimage
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-eqiad
* 11:25 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-eqiad
* 11:25 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-eqiad
* 11:24 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-eqiad
* 11:24 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-eqiad
* 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-eqiad
* 11:23 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-eqiad
* 11:23 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-eqiad
* 11:12 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db2235.codfw.wmnet with OS trixie
* 11:11 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:11 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[2160,2235].codfw.wmnet with reason: Upgrading
* 10:56 root@cumin1003: START - Cookbook sre.mysql.pool pool db2186: Maintenance
* 10:54 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1037: testing
* 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1037: testing
* 10:53 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1037: testing
* 10:53 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1037: testing
* 10:52 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1038: testing
* 10:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2186 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95072 and previous config saved to /var/cache/conftool/dbconfig/20260723-104956-cwilliams.json
* 10:49 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2186.codfw.wmnet with reason: Maintenance
* 10:43 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1054: testing
* 10:41 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1070.eqiad.wmnet with OS trixie
* 10:30 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1037 hosts
* 10:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 10:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 10:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage
* 10:16 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1070.eqiad.wmnet with reason: host reimage
* 10:06 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1038: testing
* 10:05 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1038: testing
* 10:05 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1038: testing
* 10:02 marostegui@dns1004: END - running authdns-update
* 10:00 marostegui@dns1004: START - running authdns-update
* 09:58 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: testing
* 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1054: testing
* 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1070
* 09:57 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1070
* 09:57 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1054: testing
* 09:57 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1054: testing
* 09:56 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1055: testing
* 09:56 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1070
* 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:56 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1070.eqiad.wmnet 165.48.64.10.in-addr.arpa 5.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:56 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003"
* 09:56 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1070 - jiji@cumin1003"
* 09:47 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1035 hosts
* 09:45 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 09:42 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1070
* 09:42 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1070.eqiad.wmnet with OS trixie
* 09:42 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1070.eqiad.wmnet
* 09:41 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1070.eqiad.wmnet
* 09:41 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1070.eqiad.wmnet
* 09:27 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es2051: testing
* 09:12 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: testing
* 09:12 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es2051: testing
* 09:11 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1055: testing
* 09:09 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:09 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1055: testing
* 09:09 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1055: testing
* 08:50 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:50 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:50 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:49 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:49 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:49 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:46 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1069.eqiad.wmnet
* 08:46 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1069.eqiad.wmnet
* 08:46 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1069.eqiad.wmnet
* 08:39 jiji@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 08:38 jiji@deploy1003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 08:38 jiji@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 08:37 jiji@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 08:35 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:10 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1069.eqiad.wmnet with OS trixie
* 07:49 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage
* 07:45 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1069.eqiad.wmnet with reason: host reimage
* 07:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1035 hosts
* 07:33 arthurtaylor@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 07:33 arthurtaylor@deploy1003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 07:32 arthurtaylor@deploy1003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1069
* 07:29 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1069
* 07:26 jiji@deploy1003: Finished scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules (duration: 06m 01s)
* 07:25 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1069
* 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:25 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1069.eqiad.wmnet 164.48.64.10.in-addr.arpa 4.6.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:25 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003"
* 07:25 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1069 - jiji@cumin1003"
* 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 07:25 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 07:24 jiji@deploy1003: jiji: Continuing with deployment
* 07:22 jiji@deploy1003: jiji: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:21 jiji@deploy1003: Started scap sync-world: {{Gerrit|1314018}} mediawiki: bump mcrouter and mesh modules
* 07:21 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s7
* 07:20 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1031.eqiad.wmnet,service=s2
* 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s7
* 07:20 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1031.eqiad.wmnet,service=s2
* 07:20 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts
* 07:20 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts
* 07:19 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 07:19 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1069
* 07:19 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1069.eqiad.wmnet with OS trixie
* 07:19 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1069.eqiad.wmnet
* 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 45 hosts
* 07:17 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1069.eqiad.wmnet
* 07:17 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1069.eqiad.wmnet
* 07:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 45 hosts
* 07:13 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 06:16 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2007 after successful Bookworm reimage, data transfer, and postflight validation; wdqs1013 also passed postflights and is enabled in conftool, but remains out of IPVS pending a rolling pybal restart to clear its stale pre-VLAN-move address
* 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2007.codfw.wmnet
* 05:58 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1013.eqiad.wmnet
* 05:54 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore scholarly data after Bookworm reimage) xfer scholarly_articles from wdqs2024.codfw.wmnet -> wdqs2016.codfw.wmnet, repooling source-only afterwards
== 2026-07-22 ==
* 23:34 eevans@cumin1003: END (PASS) - Cookbook sre.cassandra.roll-restart (exit_code=0) for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003
* 23:14 eevans@cumin1003: START - Cookbook sre.cassandra.roll-restart for nodes matching A:cassandra-dev: Apply upgrade to JVM17 - eevans@cumin1003
* 22:06 ryankemper: [WDQS] Added requestctl per-IP ratelimit `wdqs_heavy_sparql_bots_jul_2026_ratelimit` (chronic heavy-query bot tier driving deadlock-remediation restarts); pruned superseded `wdqs_2026_05_11_worobot`
* 21:51 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] (duration: 11m 52s)
* 21:44 sbassett@deploy1003: sbassett: Continuing with deployment
* 21:43 sbassett@deploy1003: sbassett: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:39 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1314116{{!}}OATHManage: Don't create recovery codes when viewing the page (T432915 T432916)]]
* 20:38 dani@deploy1003: Finished scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] (duration: 32m 51s)
* 20:38 ryankemper: [WDQS] Pruned obsolete requestctl action+pattern `wdqs_20260715_p2003_ring_ja3n` (actor rotated JA3Ns; rule inert)
* 20:26 dani@deploy1003: dani, vadymts1: Continuing with deployment
* 20:24 dani@deploy1003: dani, vadymts1: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2244: Testing
* 20:06 dani@deploy1003: Started scap sync-world: Backport for [[gerrit:1314046{{!}}Deploy Referring Experiences survey on enwiki (T432289)]], [[gerrit:1314106{{!}}Set $wgAutoConfirmCount to 10 for enwikiquote (T432895)]]
* 19:56 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2016.codfw.wmnet with OS bookworm
* 19:40 mutante: gerrit - one more service restart is needed - restarting
* 19:29 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2244: Testing
* 19:27 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2244: Testing
* 19:27 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2244: Testing
* 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage
* 19:15 dancy@deploy1003: Finished deploy [zuul/deploy@d92e238]: Freshening Zuul installation (duration: 00m 15s)
* 19:14 dancy@deploy1003: Started deploy [zuul/deploy@d92e238]: Freshening Zuul installation
* 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2016.codfw.wmnet with reason: host reimage
* 18:54 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1013.eqiad.wmnet, repooling source-only afterwards
* 18:52 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2016
* 18:51 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2016
* 18:51 dancy@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 18:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2016.codfw.wmnet with OS bookworm
* 18:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 11s)
* 18:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 18:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 18:30 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d] (duration: 02m 09s)
* 18:28 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (thin): Regular analytics weekly train THIN [analytics/refinery@2a25417d]
* 18:28 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d] (duration: 04m 31s)
* 18:27 dduvall: deploying https://gerrit.wikimedia.org/r/c/integration/config/+/1314025 (4 jobs updated)
* 18:23 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417]: Regular analytics weekly train [analytics/refinery@2a25417d]
* 18:22 amastilovic@deploy1003: Finished deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d] (duration: 01m 59s)
* 18:20 amastilovic@deploy1003: Started deploy [analytics/refinery@2a25417] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@2a25417d]
* 17:56 Raine: deployment server switchover => deploy1003 is primary now
* 17:55 kamila@deploy1003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 28m 24s)
* 17:54 mutante: restarting gerrit for maintenance
* 17:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1013.eqiad.wmnet with OS bookworm
* 17:27 kamila@deploy1003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]]
* 17:20 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 22m 50s)
* 17:12 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards
* 17:04 Raine: point deployment.eqiad.wmnet to deploy1003
* 17:04 kamila@dns7001: END - running authdns-update
* 17:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage
* 17:02 kamila@dns7001: START - running authdns-update
* 17:01 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover
* 17:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1013.eqiad.wmnet with reason: host reimage
* 16:58 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]]
* 16:57 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]] (duration: 02m 33s)
* 16:55 kamila@deploy2003: Locking from deployment [MediaWiki]: deployment server switchover - [[phab:T423714|T423714]]
* 16:55 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]] (duration: 00m 11s)
* 16:54 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2003 - [[phab:T240266|T240266]]
* 16:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1013
* 16:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs1013
* 16:39 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs1013
* 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:39 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs1013.eqiad.wmnet 105.32.64.10.in-addr.arpa 5.0.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:39 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:39 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003"
* 16:39 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs1013 - bking@cumin2003"
* 16:34 bking@cumin2003: START - Cookbook sre.dns.netbox
* 16:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1013
* 16:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1013.eqiad.wmnet with OS bookworm
* 16:28 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1023.eqiad.wmnet -> wdqs1024.eqiad.wmnet, repooling source-only afterwards
* 16:27 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 16:27 eevans@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linked-artifacts: apply
* 16:26 eevans@deploy2003: helmfile [eqiad] START helmfile.d/services/linked-artifacts: apply
* 16:26 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply
* 16:26 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply
* 16:25 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 29s)
* 16:25 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 16:24 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 16:21 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 16:18 eevans@deploy2003: helmfile [codfw] DONE helmfile.d/services/linked-artifacts: apply
* 16:18 eevans@deploy2003: helmfile [codfw] START helmfile.d/services/linked-artifacts: apply
* 16:08 eevans@deploy2003: helmfile [staging] DONE helmfile.d/services/linked-artifacts: apply
* 16:07 eevans@deploy2003: helmfile [staging] START helmfile.d/services/linked-artifacts: apply
* 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 16:04 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 16:01 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 15:58 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 15:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1024.eqiad.wmnet with OS bookworm
* 15:49 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 10s)
* 15:49 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]]
* 15:48 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]] (duration: 00m 15s)
* 15:48 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit1003 - [[phab:T240266|T240266]]
* 15:47 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 10s)
* 15:47 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]]
* 15:46 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:42 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:42 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:40 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:37 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 15:37 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:36 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:36 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 15:33 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-superset-metrics: apply
* 15:33 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:31 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:28 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 15:27 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:27 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage
* 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host mc-gp2006
* 15:23 jhancock@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host mc-gp2006
* 15:23 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:20 jiji@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1068.eqiad.wmnet
* 15:20 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1068.eqiad.wmnet
* 15:20 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1068.eqiad.wmnet
* 15:20 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 15:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1024.eqiad.wmnet with reason: host reimage
* 15:11 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:55 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 14:52 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1024.eqiad.wmnet with OS bookworm
* 14:50 dancy@deploy2003: Finished deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]] (duration: 00m 09s)
* 14:50 dancy@deploy2003: Started deploy [gerrit/gerrit@2fa4c79]: Applying updated motd plugin on gerrit2002 - [[phab:T240266|T240266]]
* 14:49 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 14:49 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 14:49 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 14:48 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 14:45 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:45 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:44 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:44 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:43 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:43 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:41 sukhe: ipvsadm --delete-service --tcp-service 10.2.1.55:8087: lvs2014 and lvs2013
* 14:39 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:39 sukhe: ipvsadm --delete-service --tcp-service 10.2.2.55:8087: [[phab:T432445|T432445]]
* 14:38 ecarg@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:38 ecarg@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:37 ecarg@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:37 ecarg@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:36 ecarg@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:34 ecarg@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch1001.eqiad.wmnet
* 14:32 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:32 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 14:31 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 14:31 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 14:28 sukhe: sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal': lvs2013
* 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal': lvs2014
* 14:26 sukhe: sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal'
* 14:24 sukhe: restart pybal on lvs1019
* 14:24 sukhe: restart pybal on lvs1020
* 14:19 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for 1036 hosts
* 14:17 bking@cumin2003: START - Cookbook sre.dns.netbox
* 14:04 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] (duration: 09m 28s)
* 13:59 kharlan@deploy2003: dreamyjazz, kharlan: Continuing with deployment
* 13:58 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch1001.eqiad.wmnet
* 13:57 kharlan@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:55 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313939{{!}}Allow users with suppressrevision to see mw-private-personal-info (T431292)]], [[gerrit:1313949{{!}}hCaptcha: Stop looking up the viewer's block on every page view (T432518)]]
* 13:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts datahubsearch[1002-1003].eqiad.wmnet
* 13:53 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:53 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 13:52 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: datahubsearch[1002-1003].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - bking@cumin2003"
* 13:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 13:42 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] (duration: 07m 30s)
* 13:42 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:40 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: Test
* 13:38 stran@deploy2003: dragoniez, stran: Continuing with deployment
* 13:37 stran@deploy2003: dragoniez, stran: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:35 bking@cumin2003: START - Cookbook sre.hosts.decommission for hosts datahubsearch[1002-1003].eqiad.wmnet
* 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024']
* 13:35 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 13:35 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313932{{!}}jawiki: remove redundant permission for confirmed group (T410655 T432850)]]
* 13:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 13:28 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024']
* 13:26 sukhe@dns1004: END - running authdns-update
* 13:25 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 13:25 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 13:24 sukhe@dns1004: START - running authdns-update
* 13:22 stran@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] (duration: 08m 20s)
* 13:21 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 13:20 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 13:19 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 13:19 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 13:19 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 13:18 stran@deploy2003: stran: Continuing with deployment
* 13:16 stran@deploy2003: stran: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:14 stran@deploy2003: Started scap sync-world: Backport for [[gerrit:1313845{{!}}Send an exposure event in the IRS instrument (T432718)]], [[gerrit:1313842{{!}}Send an exposure event in the IRS instrument (T432718)]]
* 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 13:13 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 13:11 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 13:08 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 12:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool es2051: Test
* 12:55 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: Test
* 12:54 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool es2051: Test
* 12:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1036 hosts
* 12:41 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:40 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts cloudvirt1048.eqiad.wmnet
* 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:40 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: cloudvirt1048.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - filippo@cumin1003"
* 12:38 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:37 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 12:37 brouberol@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 12:36 brouberol@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 12:36 brouberol@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 12:36 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:36 elukey@cumin1003: DONE (PASS) - Cookbook sre.puppet.renew-cert (exit_code=0) for crm2001.codfw.wmnet: Renew puppet certificate - elukey@cumin1003
* 12:35 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 12:35 brouberol@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 12:34 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 12:31 brouberol@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 12:30 filippo@cumin1003: START - Cookbook sre.hosts.decommission for hosts cloudvirt1048.eqiad.wmnet
* 12:30 brouberol@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 12:28 brouberol@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 12:27 brouberol@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 12:20 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1068.eqiad.wmnet with OS trixie
* 12:01 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] (duration: 13m 19s)
* 11:58 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage
* 11:52 jiji@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1068.eqiad.wmnet with reason: host reimage
* 11:51 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 11:49 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:47 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313918{{!}}Disable SI special page on enwikivoyage]]
* 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Security updates
* 11:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:43 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:42 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Security updates
* 11:42 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/analytics-test: apply
* 11:40 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/analytics-test: apply
* 11:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 11:37 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2252.codfw.wmnet with OS trixie
* 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1068
* 11:34 jiji@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1068
* 11:26 jiji@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1068
* 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:26 jiji@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1068.eqiad.wmnet 46.48.64.10.in-addr.arpa 6.4.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:26 jiji@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003"
* 11:26 jiji@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1068 - jiji@cumin1003"
* 11:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2252.codfw.wmnet with reason: host reimage
* 11:17 jiji@cumin1003: START - Cookbook sre.dns.netbox
* 11:17 jiji@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1068
* 11:17 jiji@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1068.eqiad.wmnet with OS trixie
* 11:17 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2252.codfw.wmnet with reason: host reimage
* 11:15 jiji@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1068.eqiad.wmnet
* 11:15 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:15 jiji@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1068.eqiad.wmnet
* 11:15 jiji@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1068.eqiad.wmnet
* 11:14 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:12 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] (duration: 11m 05s)
* 11:11 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:10 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:07 dreamyjazz@deploy2003: dreamyjazz, kharlan: Continuing with deployment
* 11:03 dreamyjazz@deploy2003: dreamyjazz, kharlan: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:03 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2252.codfw.wmnet with OS trixie
* 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2252: Upgrading db2252.codfw.wmnet
* 11:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:02 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2252: Upgrading db2252.codfw.wmnet
* 11:02 cwilliams@cumin1003: dbmaint on ms3@codfw [[phab:T432321|T432321]]
* 11:01 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 11:01 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1153.eqiad.wmnet with reason: Security updates
* 11:01 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1307751{{!}}WikimediaAntiAbuse: Enable everywhere (T431023)]], [[gerrit:1313187{{!}}WikimediaAntiAbuse: Enable personal info tagging on testwiki (T431292)]]
* 11:00 fnegri@deploy2003: helmfile [eqiad] DONE helmfile.d/services/toolhub: apply
* 10:58 fnegri@deploy2003: helmfile [eqiad] START helmfile.d/services/toolhub: apply
* 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1151: Security updates
* 10:57 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:57 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:57 root@cumin1003: START - Cookbook sre.mysql.pool pool db1151: Security updates
* 10:55 fnegri@deploy2003: helmfile [codfw] DONE helmfile.d/services/toolhub: apply
* 10:54 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] (duration: 08m 38s)
* 10:53 fnegri@deploy2003: helmfile [codfw] START helmfile.d/services/toolhub: apply
* 10:53 fnegri@deploy2003: helmfile [staging] DONE helmfile.d/services/toolhub: apply
* 10:52 fnegri@deploy2003: helmfile [staging] START helmfile.d/services/toolhub: apply
* 10:50 zabe@deploy2003: zabe: Continuing with deployment
* 10:47 zabe@deploy2003: zabe: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1313339{{!}}Update interwiki cache (T429921)]]
* 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1151: Security updates
* 10:42 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:42 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:42 root@cumin1003: START - Cookbook sre.mysql.depool depool db1151: Security updates
* 10:38 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] (duration: 12m 47s)
* 10:34 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 10:34 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 10:33 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2253.codfw.wmnet with OS trixie
* 10:28 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:26 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313885{{!}}Display SI in svwiki and enwikivoyage]]
* 10:18 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2253.codfw.wmnet with reason: host reimage
* 10:13 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2253.codfw.wmnet with reason: host reimage
* 10:00 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2253.codfw.wmnet with OS trixie
* 09:58 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db1151.eqiad.wmnet with reason: Security updates
* 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2253: Upgrading db2253.codfw.wmnet
* 09:57 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:57 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2253: Upgrading db2253.codfw.wmnet
* 09:56 cwilliams@cumin1003: dbmaint on ms2@codfw [[phab:T432321|T432321]]
* 09:56 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003"
* 09:36 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003
* 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import: UI improvement; support url shortener - oblivian@cumin1003
* 09:35 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import: UI improvement; support url shortener - oblivian@cumin1003"
* 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1152: Security updates
* 09:26 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:26 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:26 root@cumin1003: START - Cookbook sre.mysql.pool pool db1152: Security updates
* 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: Security updates
* 09:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:11 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:11 root@cumin1003: START - Cookbook sre.mysql.depool depool db1152: Security updates
* 09:10 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1018.eqiad.wmnet with reason: Cloning
* 09:09 Dreamy_Jazz: Deployed patch for [[phab:T432453|T432453]] and [[phab:T432454|T432454]]
* 09:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 09:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2251.codfw.wmnet with OS trixie
* 08:50 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2251.codfw.wmnet with reason: host reimage
* 08:45 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2251.codfw.wmnet with reason: host reimage
* 08:40 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] (duration: 12m 26s)
* 08:38 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1030.eqiad.wmnet,service=s1
* 08:36 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 08:31 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2251.codfw.wmnet with OS trixie
* 08:30 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:28 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313832{{!}}CheckUser SI: Enable on ukwiki and enwikivoyage without UI]]
* 08:25 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] (duration: 07m 59s)
* 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: Upgrading db2251.codfw.wmnet
* 08:22 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:22 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2251: Upgrading db2251.codfw.wmnet
* 08:20 urbanecm@deploy2003: urbanecm: Continuing with deployment
* 08:20 cwilliams@cumin1003: dbmaint on ms1@codfw [[phab:T432321|T432321]]
* 08:20 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 08:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1313823{{!}}[Growth] Set revise tone threshold to 0.79 (T432790)]]
* 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade
* 08:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade
* 08:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on db2251.codfw.wmnet,db1152.eqiad.wmnet with reason: OS upgrade
* 08:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade
* 08:13 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade
* 08:11 Dreamy_Jazz: Created cusi_signal, cusi_case, and cusi_user on ukwiki and enwikivoyage in extension1
* 08:11 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1152: OS upgrade
* 08:11 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1152: OS upgrade
* 08:04 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1030.eqiad.wmnet,service=s1
* 08:04 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1030.eqiad.wmnet,service=s1
* 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wmde: apply
* 07:53 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 07:52 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:51 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-sre: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-research: apply
* 07:50 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-research: apply
* 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-platform-eng: apply
* 07:49 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 07:48 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:47 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 07:47 phuedx: End of UTC morning backport window
* 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:43 phuedx@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] (duration: 13m 44s)
* 07:43 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 07:39 phuedx@deploy2003: phuedx: Continuing with deployment
* 07:31 phuedx@deploy2003: phuedx: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:29 phuedx@deploy2003: Started scap sync-world: Backport for [[gerrit:1313554{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]], [[gerrit:1313553{{!}}Bump VEFU schema and remove hCaptcha AB test (T432737)]]
* 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003"
* 07:24 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003
* 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Turnilo import support - oblivian@cumin1003
* 07:23 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Turnilo import support - oblivian@cumin1003"
* 06:42 ryankemper: [WDQS] [[phab:T430880|T430880]] Repooled wdqs2020 after successful Bookworm reimage, data transfer, and postflight validation
* 06:42 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet
* 05:32 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis bolwiki in section s5
* 05:25 marostegui@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis bolwiki in section s5
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-21 ==
* 22:50 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards
* 22:47 cwhite: force reboot arclamp2001 - appears to have run out of memory and gone unresponsive
* 22:24 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 01m 26s)
* 22:24 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 22:23 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 22:22 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024']
* 22:11 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 21:54 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024']
* 21:54 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 21:53 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs1024']
* 21:53 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 21:49 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2019.codfw.wmnet -> wdqs2020.codfw.wmnet, repooling source-only afterwards
* 20:57 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] (duration: 09m 10s)
* 20:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm
* 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2020.codfw.wmnet with OS bookworm
* 20:52 krinkle@deploy2003: krinkle: Continuing with deployment
* 20:49 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:47 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313340{{!}}Fixup closure setting $wgMathInternalRestbaseURL]]
* 20:45 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] (duration: 05m 42s)
* 20:44 krinkle@deploy2003: krinkle: Rolling back deployment
* 20:41 krinkle@deploy2003: krinkle: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:39 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1313235{{!}}Set $wgMathInternalRestbaseURL explicitly (take 3) (T349582)]]
* 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2003.codfw.wmnet
* 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1003.eqiad.wmnet
* 20:33 dani@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] (duration: 11m 15s)
* 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2003.codfw.wmnet
* 20:33 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1003.eqiad.wmnet
* 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage
* 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1002.eqiad.wmnet
* 20:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2002.codfw.wmnet
* 20:30 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 20:29 dani@deploy2003: dani: Continuing with deployment
* 20:29 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2020.codfw.wmnet with reason: host reimage
* 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1002.eqiad.wmnet
* 20:26 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2002.codfw.wmnet
* 20:24 dani@deploy2003: dani: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk2001.codfw.wmnet
* 20:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host flink-zk1001.eqiad.wmnet
* 20:22 dani@deploy2003: Started scap sync-world: Backport for [[gerrit:1313323{{!}}Pre-Deploy Referring Experiences survey on enwiki (T432289)]]
* 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk2001.codfw.wmnet
* 20:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host flink-zk1001.eqiad.wmnet
* 20:14 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] (duration: 09m 01s)
* 20:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm
* 20:10 sbisson@deploy2003: sbisson: Continuing with deployment
* 20:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1313200{{!}}Restore the server-side SPARQL endpoint config for Wikidata (T421250)]], [[gerrit:1311489{{!}}Article Guidance: migrate wikidata config (T421250)]]
* 20:03 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] (duration: 07m 04s)
* 20:01 mutante: Gerrit - tomorrow a new SSH host key will appear - it will be {{Gerrit|ed25519}} and has already been added to wmf-laptop. you can verify it here: https://wikitech.wikimedia.org/wiki/Help:SSH_Fingerprints/gerrit.wikimedia.org:29418 ([[phab:T240266|T240266]])
* 19:59 zabe@deploy2003: zabe: Continuing with deployment
* 19:58 zabe@deploy2003: zabe: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:56 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312604{{!}}Activate bolwiki (T429921)]]
* 19:52 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] (duration: 07m 25s)
* 19:48 zabe@deploy2003: zabe: Continuing with deployment
* 19:47 zabe@deploy2003: zabe: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:45 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1312603{{!}}Prepare Wikipedia Bole (T429921)]]
* 19:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 19:32 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wdqs1024']
* 19:27 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 19:26 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024']
* 19:26 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs1024']
* 19:12 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled `wdqs-scholarly` discovery in `eqiad` after validating `wdqs1023` end-to-end; `wdqs1024` remains disabled pending reimage recovery
* 19:11 ryankemper: [wdqs] [[phab:T430880|T430880]] Repooled wdqs1012.eqiad.wmnet after successful Bookworm reimage, data transfer, service checks, readiness probe, and cross-graph federation query validation
* 19:10 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 19:10 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1012.eqiad.wmnet
* 19:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs1024']
* 18:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for deploy1003.eqiad.wmnet
* 18:57 kamila@cumin1003: START - Cookbook sre.hosts.remove-downtime for deploy1003.eqiad.wmnet
* 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:37 sukhe@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003"
* 18:37 sukhe@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: adding urldownloader service IPs - sukhe@cumin1003"
* 18:32 sukhe@cumin1003: START - Cookbook sre.dns.netbox
* 18:32 dancy@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 18:30 sukhe@cumin1003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 18:27 sukhe@cumin1003: START - Cookbook sre.dns.netbox
* 18:24 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1024.eqiad.wmnet with OS bookworm
* 18:20 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host deploy1003.eqiad.wmnet with OS bookworm
* 18:09 kamila@deploy2003: Unlocked for deployment [MediaWiki]: deploy1003 reimage (duration: 121m 16s)
* 18:03 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 18:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1180: Security updates
* 17:55 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs2003.codfw.wmnet
* 17:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs2003.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1155.eqiad.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1155.eqiad.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2224.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2224.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2217.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2217.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2193.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2193.codfw.wmnet
* 17:38 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2180.codfw.wmnet
* 17:38 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2180.codfw.wmnet
* 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1168.eqiad.wmnet
* 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1168.eqiad.wmnet
* 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2169.codfw.wmnet
* 17:36 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2169.codfw.wmnet
* 17:36 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1165.eqiad.wmnet
* 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1165.eqiad.wmnet
* 17:35 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2158.codfw.wmnet
* 17:35 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2158.codfw.wmnet
* 17:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wcqs1003.eqiad.wmnet
* 17:17 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1180: Security updates
* 17:16 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1180.eqiad.wmnet
* 17:16 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1180.eqiad.wmnet
* 17:15 brett@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp3073.*
* 17:13 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wcqs1003.eqiad.wmnet
* 17:11 brett@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cp3073.esams.wmnet with OS trixie
* 17:11 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 17:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1024
* 17:04 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1024
* 17:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1024.eqiad.wmnet with OS bookworm
* 17:00 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 16:59 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2242: codfw rack B7 depool for maintenance
* 16:59 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 16:43 brett@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cp3073.esams.wmnet with reason: host reimage
* 16:42 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 16:39 brett@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cp3073.esams.wmnet with reason: host reimage
* 16:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage
* 16:27 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on deploy1003.eqiad.wmnet with reason: host reimage
* 16:14 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2242: codfw rack B7 depool for maintenance
* 16:14 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: codfw rack B7 depool for maintenance
* 16:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1023.eqiad.wmnet with OS bookworm
* 16:13 brett@cumin2002: START - Cookbook sre.hosts.reimage for host cp3073.esams.wmnet with OS trixie
* 16:08 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host deploy1003.eqiad.wmnet with OS bookworm
* 16:08 kamila@deploy2003: Locking from deployment [MediaWiki]: deploy1003 reimage
* 16:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 16:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1012.eqiad.wmnet with OS bookworm
* 15:48 inflatador: bking@apt1002 `sudo reprepro copy bookworm-wikimedia bullseye-wikimedia jvmquake` [[phab:T430880|T430880]]
* 15:39 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp3073.*
* 15:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 15:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:34 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1311536{{!}}Set $wgMathInternalRestbaseURL explicitly (take 2) (T349582)]]
* 15:29 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:29 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2228: codfw rack B7 depool for maintenance
* 15:29 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: codfw rack B7 depool for maintenance
* 15:27 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1180: Security update
* 15:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update
* 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 15:21 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 15:19 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 24s)
* 15:19 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 15:14 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-dumps: apply
* 15:14 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.depool (exit_code=99) depool db1180: Security update
* 15:13 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-dumps: apply
* 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-eqiad
* 14:48 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-eqiad
* 14:44 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool db2229: codfw rack B7 depool for maintenance
* 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc2017: codfw rack B7 depool for maintenance
* 14:44 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache
* 14:43 cmooney@cumin2003: START - Cookbook sre.mysql.pool pool pc2017: codfw rack B7 depool for maintenance
* 14:43 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-staging2003.codfw.wmnet
* 14:43 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-staging2003.codfw.wmnet
* 14:42 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2009.codfw.wmnet
* 14:42 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2009.codfw.wmnet
* 14:41 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:41 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:40 cmooney@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 29 hosts
* 14:40 cmooney@cumin1003: START - Cookbook sre.hosts.remove-downtime for 29 hosts
* 14:35 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 14:34 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 14:32 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] (duration: 07m 56s)
* 14:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ssw1-a[1,8]-codfw with reason: lsw1-b7-codfw JunOS upgrade
* 14:28 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 14:28 elukey@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 14:26 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:24 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1313194{{!}}CheckUser Suggested Investigations: Enable for 4 more wikis]]
* 14:23 topranks: reboot lsw1-b7-codfw to upgrade JunOS (affects all hosts in rack) [[phab:T430928|T430928]]
* 14:18 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-staging2003.codfw.wmnet
* 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2009.codfw.wmnet
* 14:14 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1180: Security update
* 14:13 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2242: codfw rack B7 depool for maintenance
* 14:13 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2242: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2228: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2229: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool db2229: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2017: codfw rack B7 depool for maintenance
* 14:12 cmooney@cumin2003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.parsercache
* 14:12 cmooney@cumin2003: START - Cookbook sre.mysql.depool depool pc2017: codfw rack B7 depool for maintenance
* 14:08 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-staging2003.codfw.wmnet
* 14:07 cmooney@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on aux-k8s-etcd2004.codfw.wmnet,ml-etcd2001.codfw.wmnet with reason: lsw1-b7-codfw JunOS upgrade
* 14:06 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95005 and previous config saved to /var/cache/conftool/dbconfig/20260721-140620-cwilliams.json
* 14:05 cmooney@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2049.codfw.wmnet
* 14:05 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-scholarly,name=eqiad
* 14:04 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2009.codfw.wmnet
* 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 14:04 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dumps: apply
* 14:03 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:03 cmooney@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-ctrl2001.codfw.wmnet
* 14:03 cmooney@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-ctrl2001.codfw.wmnet
* 14:02 cmooney@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2049.codfw.wmnet
* 14:00 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2139-2140,2157,2284-2285].codfw.wmnet
* 14:00 Dreamy_Jazz: Created cusi_case, cusi_signal, and cusi_user on svwiki, dewiki, jawiki, eswiki
* 13:59 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b7-codfw,lsw1-b7-codfw IPv6,lsw1-b7-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: lsw1-b7-codfw JunOS upgrade
* 13:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards
* 13:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 29 hosts with reason: lsw1-b7-codfw JunOS upgrade
* 13:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95003 and previous config saved to /var/cache/conftool/dbconfig/20260721-135613-cwilliams.json
* 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage
* 13:53 cmooney@cumin1003: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 1:00:00 on 30 hosts with reason: lsw1-b7-codfw JunOS upgrade
* 13:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1012.eqiad.wmnet with reason: host reimage
* 13:48 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 13:48 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 13:46 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224', diff saved to https://phabricator.wikimedia.org/P95001 and previous config saved to /var/cache/conftool/dbconfig/20260721-134605-cwilliams.json
* 13:46 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet
* 13:46 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet
* 13:45 elukey: move the Docker Registry's /v2/wikimedia/machinelearning.* prefix to the ml S3 backend - [[phab:T428022|T428022]]
* 13:45 cmooney@cumin1003: END (FAIL) - Cookbook sre.ganeti.drain-node (exit_code=99) for draining ganeti node ganeti2033.codfw.wmnet
* 13:45 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet
* 13:43 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:40 jiji@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-mcrouter: apply
* 13:40 jiji@deploy2003: helmfile [codfw] START helmfile.d/services/mw-mcrouter: apply
* 13:39 jiji@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-mcrouter: apply
* 13:39 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 13:38 cmooney@cumin1003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2032.codfw.wmnet
* 13:38 jiji@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-mcrouter: apply
* 13:38 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 13:37 cmooney@cumin1003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2032.codfw.wmnet
* 13:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P95000 and previous config saved to /var/cache/conftool/dbconfig/20260721-133557-cwilliams.json
* 13:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1012
* 13:33 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1012
* 13:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1012.eqiad.wmnet with OS bookworm
* 13:30 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:30 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2224 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94999 and previous config saved to /var/cache/conftool/dbconfig/20260721-132855-cwilliams.json
* 13:28 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2224.codfw.wmnet with reason: Maintenance
* 13:28 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94998 and previous config saved to /var/cache/conftool/dbconfig/20260721-132826-cwilliams.json
* 13:28 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 13:23 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] (duration: 07m 50s)
* 13:20 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 13:18 kharlan@deploy2003: kharlan: Continuing with deployment
* 13:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94996 and previous config saved to /var/cache/conftool/dbconfig/20260721-131817-cwilliams.json
* 13:17 brouberol@dns1004: END - running authdns-update
* 13:17 kharlan@deploy2003: kharlan: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:15 brouberol@dns1004: START - running authdns-update
* 13:15 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1313092{{!}}diff: Mark UnifiedDiffFormatter as stable to extend (T432457)]], [[gerrit:1313093{{!}}Add RevisionSnippetGenerator service (T432457)]]
* 13:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94995 and previous config saved to /var/cache/conftool/dbconfig/20260721-131411-cwilliams.json
* 13:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs1024.eqiad.wmnet -> wdqs1023.eqiad.wmnet, repooling source-only afterwards
* 13:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217', diff saved to https://phabricator.wikimedia.org/P94994 and previous config saved to /var/cache/conftool/dbconfig/20260721-130809-cwilliams.json
* 13:07 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 13:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94993 and previous config saved to /var/cache/conftool/dbconfig/20260721-130404-cwilliams.json
* 13:03 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:03 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:02 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 13:02 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 12:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94992 and previous config saved to /var/cache/conftool/dbconfig/20260721-125801-cwilliams.json
* 12:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180', diff saved to https://phabricator.wikimedia.org/P94991 and previous config saved to /var/cache/conftool/dbconfig/20260721-125356-cwilliams.json
* 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2217 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94990 and previous config saved to /var/cache/conftool/dbconfig/20260721-125049-cwilliams.json
* 12:50 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2217.codfw.wmnet with reason: Maintenance
* 12:50 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94989 and previous config saved to /var/cache/conftool/dbconfig/20260721-125017-cwilliams.json
* 12:48 elukey: bmc cold reboot for lvs1013 and lvs1015 - [[phab:T426180|T426180]]
* 12:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94988 and previous config saved to /var/cache/conftool/dbconfig/20260721-124348-cwilliams.json
* 12:40 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94987 and previous config saved to /var/cache/conftool/dbconfig/20260721-124009-cwilliams.json
* 12:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts
* 12:33 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts
* 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 2 hosts
* 12:32 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 2 hosts
* 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:30 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193', diff saved to https://phabricator.wikimedia.org/P94986 and previous config saved to /var/cache/conftool/dbconfig/20260721-123001-cwilliams.json
* 12:19 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94985 and previous config saved to /var/cache/conftool/dbconfig/20260721-121953-cwilliams.json
* 12:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2007.codfw.wmnet with OS bookworm
* 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2193 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94983 and previous config saved to /var/cache/conftool/dbconfig/20260721-121257-cwilliams.json
* 12:12 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2193.codfw.wmnet with reason: Maintenance
* 12:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94982 and previous config saved to /var/cache/conftool/dbconfig/20260721-121239-cwilliams.json
* 12:02 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94980 and previous config saved to /var/cache/conftool/dbconfig/20260721-120231-cwilliams.json
* 11:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180', diff saved to https://phabricator.wikimedia.org/P94979 and previous config saved to /var/cache/conftool/dbconfig/20260721-115223-cwilliams.json
* 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94978 and previous config saved to /var/cache/conftool/dbconfig/20260721-114333-cwilliams.json
* 11:43 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1180.eqiad.wmnet with reason: Maintenance
* 11:43 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94977 and previous config saved to /var/cache/conftool/dbconfig/20260721-114305-cwilliams.json
* 11:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94976 and previous config saved to /var/cache/conftool/dbconfig/20260721-114215-cwilliams.json
* 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2180 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94975 and previous config saved to /var/cache/conftool/dbconfig/20260721-113530-cwilliams.json
* 11:35 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2180.codfw.wmnet with reason: Maintenance
* 11:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94974 and previous config saved to /var/cache/conftool/dbconfig/20260721-113501-cwilliams.json
* 11:32 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94973 and previous config saved to /var/cache/conftool/dbconfig/20260721-113258-cwilliams.json
* 11:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94972 and previous config saved to /var/cache/conftool/dbconfig/20260721-112453-cwilliams.json
* 11:22 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168', diff saved to https://phabricator.wikimedia.org/P94971 and previous config saved to /var/cache/conftool/dbconfig/20260721-112250-cwilliams.json
* 11:21 XioNoX: put eqiad-drmrs Arelion link in service
* 11:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169', diff saved to https://phabricator.wikimedia.org/P94970 and previous config saved to /var/cache/conftool/dbconfig/20260721-111446-cwilliams.json
* 11:12 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94969 and previous config saved to /var/cache/conftool/dbconfig/20260721-111242-cwilliams.json
* 11:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 11:10 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 11:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1093 hosts
* 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1168 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94968 and previous config saved to /var/cache/conftool/dbconfig/20260721-110548-cwilliams.json
* 11:05 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1168.eqiad.wmnet with reason: Maintenance
* 11:05 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94967 and previous config saved to /var/cache/conftool/dbconfig/20260721-110520-cwilliams.json
* 11:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94966 and previous config saved to /var/cache/conftool/dbconfig/20260721-110439-cwilliams.json
* 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2169 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94964 and previous config saved to /var/cache/conftool/dbconfig/20260721-105632-cwilliams.json
* 10:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2169.codfw.wmnet with reason: Maintenance
* 10:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94963 and previous config saved to /var/cache/conftool/dbconfig/20260721-105603-cwilliams.json
* 10:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94962 and previous config saved to /var/cache/conftool/dbconfig/20260721-105512-cwilliams.json
* 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94961 and previous config saved to /var/cache/conftool/dbconfig/20260721-104555-cwilliams.json
* 10:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165', diff saved to https://phabricator.wikimedia.org/P94960 and previous config saved to /var/cache/conftool/dbconfig/20260721-104504-cwilliams.json
* 10:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158', diff saved to https://phabricator.wikimedia.org/P94959 and previous config saved to /var/cache/conftool/dbconfig/20260721-103547-cwilliams.json
* 10:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94958 and previous config saved to /var/cache/conftool/dbconfig/20260721-103456-cwilliams.json
* 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2229: Upgraded kernel
* 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1165 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94956 and previous config saved to /var/cache/conftool/dbconfig/20260721-102757-cwilliams.json
* 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on an-redacteddb1001.eqiad.wmnet,clouddb[1015,1025,1028].eqiad.wmnet,db1155.eqiad.wmnet with reason: Maintenance
* 10:27 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1165.eqiad.wmnet with reason: Maintenance
* 10:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94955 and previous config saved to /var/cache/conftool/dbconfig/20260721-102539-cwilliams.json
* 10:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2158 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94954 and previous config saved to /var/cache/conftool/dbconfig/20260721-101848-cwilliams.json
* 10:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2158.codfw.wmnet with reason: Maintenance
* 09:43 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2229: Upgraded kernel
* 09:42 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2229.codfw.wmnet
* 09:42 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2229.codfw.wmnet
* 09:23 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db2229.codfw.wmnet
* 09:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2229.codfw.wmnet
* 08:57 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2229 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94948 and previous config saved to /var/cache/conftool/dbconfig/20260721-085724-cwilliams.json
* 08:54 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2214 to s6 primary [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94947 and previous config saved to /var/cache/conftool/dbconfig/20260721-085442-cwilliams.json
* 08:53 cezmunsta: Starting s6 codfw failover from db2229 to db2214 - [[phab:T430964|T430964]]
* 08:46 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2214 with weight 0 [[phab:T430964|T430964]]', diff saved to https://phabricator.wikimedia.org/P94946 and previous config saved to /var/cache/conftool/dbconfig/20260721-084613-cwilliams.json
* 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 22 hosts with reason: Primary switchover s6 [[phab:T430964|T430964]]
* 08:32 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1017.eqiad.wmnet,service=s1
* 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003
* 08:06 ayounsi@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Add subrated circuit rate to interface descriptions - CR1312476 - ayounsi@cumin1003
* 07:58 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] (duration: 12m 55s)
* 07:51 reedy@deploy2003: reedy, neriah: Continuing with deployment
* 07:51 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs-next: apply
* 07:51 reedy@deploy2003: reedy, neriah: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1093 hosts
* 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm2001.wikimedia.org
* 07:45 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1312614{{!}}Enable UserPageEditProtection on jawiki (T392754 T410655)]]
* 07:43 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs-next: apply
* 07:42 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm2001.wikimedia.org
* 07:23 elukey: upgrade libtiff6 packages on zuul* trixie hosts for security upgrades
* 07:22 elukey: upgrade libtiff6 packages on Wikikube trixie workers for security upgrades
* 07:14 elukey@deploy2003: helmfile [codfw] DONE helmfile.d/services/proton: sync
* 07:13 elukey@deploy2003: helmfile [codfw] START helmfile.d/services/proton: sync
* 07:11 elukey@deploy2003: helmfile [eqiad] DONE helmfile.d/services/proton: sync
* 07:10 elukey@deploy2003: helmfile [eqiad] START helmfile.d/services/proton: sync
* 07:09 elukey@deploy2003: helmfile [staging] DONE helmfile.d/services/proton: sync
* 07:08 elukey@deploy2003: helmfile [staging] START helmfile.d/services/proton: sync
* 06:54 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage
* 06:46 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1023.eqiad.wmnet with reason: host reimage
* 06:24 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003"
* 05:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003
* 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Haproxy-only mode support - oblivian@cumin1003
* 05:42 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Haproxy-only mode support - oblivian@cumin1003"
* 05:38 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1017.eqiad.wmnet with reason: Cloning
* 05:37 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1017.eqiad.wmnet,service=s1
* 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s8
* 05:33 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1029.eqiad.wmnet,service=s5
* 05:32 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:30 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 05:11 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1003.eqiad.wmnet
* 05:04 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1003.eqiad.wmnet
* 05:00 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 04:56 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.9 (duration: 01m 08s)
* 03:41 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]] (duration: 36m 30s)
* 03:05 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.12 refs [[phab:T430831|T430831]]
* 03:01 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 03:01 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 03:00 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 03:00 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:45 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:44 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:43 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:41 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:36 ssastry@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 02:36 ssastry@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 02:36 ssastry@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 02:35 ssastry@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 02:16 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 47s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 00:56 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
== 2026-07-20 ==
* 23:38 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 23:07 Amir1: deleting echo notifications from 2015 on group1 wikis
* 23:07 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] (duration: 14m 16s)
* 23:01 ladsgroup@deploy2003: ladsgroup: Continuing with deployment
* 23:00 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1312302{{!}}Disable MJPEG, enable MPEG-4 Part II (T358266)]]
* 22:46 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards
* 22:39 maryum: Deployed security fixes for several security bugs
* 21:42 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 21:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2008.codfw.wmnet -> wdqs2007.codfw.wmnet, repooling source-only afterwards
* 21:37 ryankemper@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 21:37 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 21:34 sbassett: Deployed security fix for [[phab:T432424|T432424]]
* 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet
* 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet
* 21:33 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet
* 21:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 17s)
* 21:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 21:27 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage
* 21:08 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-main,name=codfw
* 21:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2007.codfw.wmnet with reason: host reimage
* 20:59 sukhe: sukhe@lvs2013:~$ sudo systemctl restart pybal.service
* 20:58 sukhe: pybal restart for IP changes around wdqs-main hosts
* 20:57 sukhe: sukhe@lvs2014:~$ sudo systemctl restart pybal.service
* 20:46 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-main,name=codfw
* 20:45 ebernhardson@deploy2003: Finished deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat (duration: 00m 34s)
* 20:44 ebernhardson@deploy2003: Started deploy [search/mjolnir/deploy@d4dc3b8]: Update for opensearch 2.x compat
* 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2007
* 20:44 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2007
* 20:43 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2007
* 20:43 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2007.codfw.wmnet 156.16.192.10.in-addr.arpa 6.5.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003"
* 20:41 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2007 - bking@cumin2003"
* 20:33 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94944 and previous config saved to /var/cache/conftool/dbconfig/20260720-203333-cwilliams.json
* 20:32 arlolra@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] (duration: 15m 07s)
* 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs2020.codfw.wmnet
* 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1023.eqiad.wmnet
* 20:30 ryankemper@cumin2003: conftool action : set/pooled=yes; selector: name=wdqs1011.eqiad.wmnet
* 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs2020.codfw.wmnet
* 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1023.eqiad.wmnet
* 20:29 ryankemper@cumin2003: conftool action : set/pooled=no; selector: name=wdqs1011.eqiad.wmnet
* 20:25 arlolra@deploy2003: arlolra, cscott: Continuing with deployment
* 20:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94943 and previous config saved to /var/cache/conftool/dbconfig/20260720-202325-cwilliams.json
* 20:21 arlolra@deploy2003: arlolra, cscott: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:17 arlolra@deploy2003: Started scap sync-world: Backport for [[gerrit:1312547{{!}}Enable PRV on enwiki talk namespace, template namespace]], [[gerrit:1312554{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T221028 T226289 T353697 T430854 T431771)]], [[gerrit:1312556{{!}}Bump wikimedia/parsoid to 0.24.0-a15 (T431771)]]
* 20:13 bking@cumin2003: START - Cookbook sre.dns.netbox
* 20:13 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257', diff saved to https://phabricator.wikimedia.org/P94942 and previous config saved to /var/cache/conftool/dbconfig/20260720-201318-cwilliams.json
* 20:13 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 20:10 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 20:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2007
* 20:04 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2007.codfw.wmnet with OS bookworm
* 20:03 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94941 and previous config saved to /var/cache/conftool/dbconfig/20260720-200310-cwilliams.json
* 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1257 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94940 and previous config saved to /var/cache/conftool/dbconfig/20260720-195633-cwilliams.json
* 19:56 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1257.eqiad.wmnet with reason: Maintenance
* 19:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94939 and previous config saved to /var/cache/conftool/dbconfig/20260720-195605-cwilliams.json
* 19:51 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 19:50 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 19:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94938 and previous config saved to /var/cache/conftool/dbconfig/20260720-194558-cwilliams.json
* 19:44 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1012.eqiad.wmnet -> wdqs1011.eqiad.wmnet, repooling source-only afterwards
* 19:43 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 19:42 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wikidata from wdqs1011.eqiad.wmnet -> wdqs1012.eqiad.wmnet, repooling source-only afterwards
* 19:41 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wdqs1011.eqiad.wmnet with OS bookworm
* 19:35 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256', diff saved to https://phabricator.wikimedia.org/P94937 and previous config saved to /var/cache/conftool/dbconfig/20260720-193550-cwilliams.json
* 19:25 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94936 and previous config saved to /var/cache/conftool/dbconfig/20260720-192542-cwilliams.json
* 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1256 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94935 and previous config saved to /var/cache/conftool/dbconfig/20260720-191856-cwilliams.json
* 19:18 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1256.eqiad.wmnet with reason: Maintenance
* 19:18 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94934 and previous config saved to /var/cache/conftool/dbconfig/20260720-191839-cwilliams.json
* 19:08 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94933 and previous config saved to /var/cache/conftool/dbconfig/20260720-190831-cwilliams.json
* 18:58 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255', diff saved to https://phabricator.wikimedia.org/P94932 and previous config saved to /var/cache/conftool/dbconfig/20260720-185824-cwilliams.json
* 18:50 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 18:48 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94931 and previous config saved to /var/cache/conftool/dbconfig/20260720-184816-cwilliams.json
* 18:42 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1255 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94930 and previous config saved to /var/cache/conftool/dbconfig/20260720-184224-cwilliams.json
* 18:42 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1255.eqiad.wmnet with reason: Maintenance
* 18:41 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94929 and previous config saved to /var/cache/conftool/dbconfig/20260720-184153-cwilliams.json
* 18:39 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 18:39 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s)
* 18:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 18:38 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 59m 26s)
* 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=99) upgrade firmware for hosts ['wdqs2020']
* 18:37 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 18:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94928 and previous config saved to /var/cache/conftool/dbconfig/20260720-183145-cwilliams.json
* 18:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211', diff saved to https://phabricator.wikimedia.org/P94927 and previous config saved to /var/cache/conftool/dbconfig/20260720-182137-cwilliams.json
* 18:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94926 and previous config saved to /var/cache/conftool/dbconfig/20260720-181129-cwilliams.json
* 18:09 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_codfw
* 18:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2057.codfw.wmnet
* 18:08 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_codfw
* 18:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2058.codfw.wmnet
* 18:04 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db1211 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94925 and previous config saved to /var/cache/conftool/dbconfig/20260720-180452-cwilliams.json
* 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb[1016,1020,1022-1023].eqiad.wmnet,db1154.eqiad.wmnet with reason: Maintenance
* 18:04 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db1211.eqiad.wmnet with reason: Maintenance
* 18:02 sukhe: armed keyholder on acmechief1002.eqiad.wmnet and acmechief2002.codfw.wmnet (active host)
* 18:01 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief2002.codfw.wmnet
* 17:57 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief2002.codfw.wmnet
* 17:56 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wdqs2020']
* 17:52 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief1002.eqiad.wmnet
* 17:50 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 17:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage
* 17:48 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief1002.eqiad.wmnet
* 17:47 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test2001.codfw.wmnet
* 17:47 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94924 and previous config saved to /var/cache/conftool/dbconfig/20260720-174717-cwilliams.json
* 17:46 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wdqs2020']
* 17:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1011.eqiad.wmnet with reason: host reimage
* 17:43 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs2020.codfw.wmnet with OS bookworm
* 17:43 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test2001.codfw.wmnet
* 17:43 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host acmechief-test1001.eqiad.wmnet
* 17:39 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host acmechief-test1001.eqiad.wmnet
* 17:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 17:38 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:38 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 17:37 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 17:37 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94923 and previous config saved to /var/cache/conftool/dbconfig/20260720-173709-cwilliams.json
* 17:35 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 17:31 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 20m 40s)
* 17:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2055.codfw.wmnet
* 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2056.codfw.wmnet
* 17:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1011
* 17:27 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1011
* 17:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1011.eqiad.wmnet with OS bookworm
* 17:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244', diff saved to https://phabricator.wikimedia.org/P94922 and previous config saved to /var/cache/conftool/dbconfig/20260720-172701-cwilliams.json
* 17:16 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94921 and previous config saved to /var/cache/conftool/dbconfig/20260720-171653-cwilliams.json
* 17:11 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 17:11 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 13m 03s)
* 17:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2244 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94920 and previous config saved to /var/cache/conftool/dbconfig/20260720-171012-cwilliams.json
* 17:10 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2244.codfw.wmnet with reason: Maintenance
* 17:09 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94919 and previous config saved to /var/cache/conftool/dbconfig/20260720-170941-cwilliams.json
* 16:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94918 and previous config saved to /var/cache/conftool/dbconfig/20260720-165933-cwilliams.json
* 16:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 16:58 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2053.codfw.wmnet
* 16:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2054.codfw.wmnet
* 16:49 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243', diff saved to https://phabricator.wikimedia.org/P94917 and previous config saved to /var/cache/conftool/dbconfig/20260720-164926-cwilliams.json
* 16:39 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94916 and previous config saved to /var/cache/conftool/dbconfig/20260720-163918-cwilliams.json
* 16:35 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wdqs1023.eqiad.wmnet with OS bookworm
* 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2243 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94915 and previous config saved to /var/cache/conftool/dbconfig/20260720-163140-cwilliams.json
* 16:31 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2243.codfw.wmnet with reason: Maintenance
* 16:31 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94914 and previous config saved to /var/cache/conftool/dbconfig/20260720-163111-cwilliams.json
* 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 16:27 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 16:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2020
* 16:23 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2020
* 16:21 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2020
* 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:21 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2020.codfw.wmnet 85.0.192.10.in-addr.arpa 5.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:21 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:21 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94913 and previous config saved to /var/cache/conftool/dbconfig/20260720-162103-cwilliams.json
* 16:19 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 16:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 16:18 bking@cumin2003: START - Cookbook sre.dns.netbox
* 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.dns.netbox (exit_code=99)
* 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:17 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002"
* 16:17 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for netbox accounting errors - jhancock@cumin2002"
* 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2051.codfw.wmnet
* 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2052.codfw.wmnet
* 16:11 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 16:10 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242', diff saved to https://phabricator.wikimedia.org/P94912 and previous config saved to /var/cache/conftool/dbconfig/20260720-161055-cwilliams.json
* 16:09 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 16:08 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 16:06 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 16:06 bking@cumin2003: START - Cookbook sre.dns.netbox
* 16:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2020
* 16:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2020.codfw.wmnet with OS bookworm
* 16:00 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94911 and previous config saved to /var/cache/conftool/dbconfig/20260720-160047-cwilliams.json
* 15:58 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards
* 15:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2242 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94909 and previous config saved to /var/cache/conftool/dbconfig/20260720-155353-cwilliams.json
* 15:53 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2242.codfw.wmnet with reason: Maintenance
* 15:44 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94908 and previous config saved to /var/cache/conftool/dbconfig/20260720-154433-cwilliams.json
* 15:35 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2049.codfw.wmnet
* 15:34 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94907 and previous config saved to /var/cache/conftool/dbconfig/20260720-153425-cwilliams.json
* 15:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2050.codfw.wmnet
* 15:28 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (3 nodes at a time) for ElasticSearch cluster search_codfw: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 15:24 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162', diff saved to https://phabricator.wikimedia.org/P94906 and previous config saved to /var/cache/conftool/dbconfig/20260720-152418-cwilliams.json
* 15:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1023
* 15:14 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1023
* 15:14 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1023.eqiad.wmnet with OS bookworm
* 15:14 cwilliams@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94905 and previous config saved to /var/cache/conftool/dbconfig/20260720-151407-cwilliams.json
* 15:13 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] (duration: 41m 16s)
* 15:08 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 15:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Depooling db2162 ([[phab:T431660|T431660]])', diff saved to https://phabricator.wikimedia.org/P94902 and previous config saved to /var/cache/conftool/dbconfig/20260720-150729-cwilliams.json
* 15:07 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2162.codfw.wmnet with reason: Maintenance
* 15:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards
* 15:00 urbanecm@deploy2003: vadymts1, migr, urbanecm: Continuing with deployment
* 14:59 bking@cumin2003: END (PASS) - Cookbook sre.elasticsearch.rolling-operation (exit_code=0) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 07s)
* 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:58 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 13s)
* 14:58 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 14:57 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2019.codfw.wmnet, repooling source-only afterwards
* 14:57 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2047.codfw.wmnet
* 14:55 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2048.codfw.wmnet
* 14:51 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2019.codfw.wmnet with OS bookworm
* 14:49 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 14:47 urbanecm@deploy2003: vadymts1, migr, urbanecm: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:44 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified]
* 14:44 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: pool magru [reason: BGP issues in lvs7003 resolved after liberica restart, no task ID specified]
* 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:41 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin pooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:41 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin depooling P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:39 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart P<nowiki>{</nowiki>lvs7003.magru.wmnet<nowiki>}</nowiki> and A:liberica
* 14:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.admin (exit_code=0) DNS admin: depool magru [reason: no reason specified, no task ID specified]
* 14:33 sukhe@cumin1003: START - Cookbook sre.dns.admin DNS admin: depool magru [reason: no reason specified, no task ID specified]
* 14:31 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1312471{{!}}postEdit experiment: enroll control users by same criteria]], [[gerrit:1312423{{!}}Modify user groups rights in English Wikiquote (T432557)]]
* 14:24 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:24 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:24 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:24 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage
* 14:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2027.codfw.wmnet, repooling source-only afterwards
* 14:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2019.codfw.wmnet with reason: host reimage
* 14:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2027.codfw.wmnet with OS bookworm
* 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2046.codfw.wmnet
* 14:16 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2045.codfw.wmnet
* 14:08 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:08 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:08 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:08 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:07 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:06 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:06 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:06 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: apply security updates - bking@cumin2003 - [[phab:T431826|T431826]]
* 14:05 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2019
* 14:00 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2019
* 13:56 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1015.eqiad.wmnet
* 13:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage
* 13:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2071.codfw.wmnet with OS trixie
* 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and not P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:51 sukhe@cumin1003: END (ERROR) - Cookbook sre.loadbalancer.admin (exit_code=97) rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:51 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting A:liberica and P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:51 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1015.eqiad.wmnet
* 13:50 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1014.eqiad.wmnet
* 13:50 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1076.eqiad.wmnet with OS trixie
* 13:50 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2027.codfw.wmnet with reason: host reimage
* 13:45 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1014.eqiad.wmnet
* 13:44 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lvs1013.eqiad.wmnet
* 13:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 13:38 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host lvs1013.eqiad.wmnet
* 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2044.codfw.wmnet
* 13:37 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp2043.codfw.wmnet
* 13:36 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2019
* 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:36 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2019.codfw.wmnet 156.32.192.10.in-addr.arpa 6.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:36 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:36 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003"
* 13:36 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2019 - bking@cumin2003"
* 13:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2005.codfw.wmnet
* 13:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage
* 13:31 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:31 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2019
* 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2005.codfw.wmnet
* 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry2004.codfw.wmnet
* 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2019.codfw.wmnet with OS bookworm
* 13:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2027
* 13:30 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2027
* 13:30 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2027.codfw.wmnet with OS bookworm
* 13:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage
* 13:29 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_codfw
* 13:28 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_codfw
* 13:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry2004.codfw.wmnet
* 13:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1005.eqiad.wmnet
* 13:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2071.codfw.wmnet with reason: host reimage
* 13:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1076.eqiad.wmnet with reason: host reimage
* 13:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1005.eqiad.wmnet
* 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host registry1004.eqiad.wmnet
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host registry1004.eqiad.wmnet
* 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts
* 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts
* 13:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts
* 13:12 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts
* 13:03 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1076.eqiad.wmnet with OS trixie
* 13:03 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2071.codfw.wmnet with OS trixie
* 12:55 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:54 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:46 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts
* 12:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts
* 12:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 7 hosts
* 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 7 hosts
* 12:42 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 3 hosts
* 12:42 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 3 hosts
* 12:40 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2070.codfw.wmnet with OS trixie
* 12:36 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:35 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1075.eqiad.wmnet with OS trixie
* 12:32 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 10 hosts
* 12:31 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 10 hosts
* 12:22 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin1001.eqiad.wmnet
* 12:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin1001.eqiad.wmnet
* 12:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage
* 12:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcumin2001.codfw.wmnet
* 12:14 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage
* 12:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2070.codfw.wmnet with reason: host reimage
* 12:10 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1075.eqiad.wmnet with reason: host reimage
* 12:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcumin2001.codfw.wmnet
* 11:17 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1074.eqiad.wmnet with OS trixie
* 11:17 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 11:16 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 11:14 ozge@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 6 hosts
* 11:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 6 hosts
* 11:06 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 324 hosts
* 10:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage
* 10:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2069.codfw.wmnet with OS trixie
* 10:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1074.eqiad.wmnet with reason: host reimage
* 10:30 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage
* 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1074.eqiad.wmnet with OS trixie
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2069.codfw.wmnet with reason: host reimage
* 10:06 blake@deploy2003: Stopping before sync operations
* 10:06 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values
* 10:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2069.codfw.wmnet with OS trixie
* 10:00 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1073.eqiad.wmnet with OS trixie
* 09:56 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 324 hosts
* 09:39 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 09:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 8 hosts
* 09:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage
* 09:37 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for 8 hosts
* 09:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1073.eqiad.wmnet with reason: host reimage
* 09:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2068.codfw.wmnet with OS trixie
* 09:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1073.eqiad.wmnet with OS trixie
* 09:13 blake@deploy2003: sync-world aborted: Non-deployment scap run to populate new release values (duration: 00m 02s)
* 09:13 blake@deploy2003: Started scap sync-world: Non-deployment scap run to populate new release values
* 08:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage
* 08:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2068.codfw.wmnet with reason: host reimage
* 08:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 08:32 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2068.codfw.wmnet with OS trixie
* 08:15 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1072.eqiad.wmnet with OS trixie
* 07:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2067.codfw.wmnet with OS trixie
* 07:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage
* 07:49 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1072.eqiad.wmnet with reason: host reimage
* 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 07:45 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 07:39 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage
* 07:35 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2067.codfw.wmnet with reason: host reimage
* 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 07:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1072.eqiad.wmnet with OS trixie
* 07:30 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 07:17 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 07:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2067.codfw.wmnet with OS trixie
* 05:51 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:50 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:25 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 05:25 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down
* 04:28 kevinbazira@deploy2003: helmfile [ml-staging-codfw] 'sync' command on namespace 'tts-section-generator' for release 'main' .
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 07m 02s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-18 ==
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 00:11 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards
== 2026-07-17 ==
* 23:53 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards
* 23:09 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs2022.codfw.wmnet -> wdqs2018.codfw.wmnet, repooling source-only afterwards
* 23:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2017.codfw.wmnet -> wdqs2026.codfw.wmnet, repooling source-only afterwards
* 22:11 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2018.codfw.wmnet with OS bookworm
* 22:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2026.codfw.wmnet with OS bookworm
* 21:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage
* 21:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2018.codfw.wmnet with reason: host reimage
* 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage
* 21:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2026.codfw.wmnet with reason: host reimage
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2018
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2018
* 21:26 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2018
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:26 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2018.codfw.wmnet 155.32.192.10.in-addr.arpa 5.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:26 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003"
* 21:26 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2018 - bking@cumin2003"
* 21:14 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:13 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2018
* 21:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2018.codfw.wmnet with OS bookworm
* 21:12 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2026
* 21:12 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2026
* 21:12 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2026.codfw.wmnet with OS bookworm
* 21:05 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards
* 20:57 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards
* 20:11 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs2017.codfw.wmnet, repooling source-only afterwards
* 20:08 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2017.codfw.wmnet with OS bookworm
* 20:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], restore data on newly-reimaged host) xfer wdqs-all from wdqs1022.eqiad.wmnet -> wdqs1026.eqiad.wmnet, repooling source-only afterwards
* 20:06 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1026.eqiad.wmnet with OS bookworm
* 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:55 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 09s)
* 19:55 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 08s)
* 19:50 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:50 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 10m 03s)
* 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage
* 19:40 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:40 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 15s)
* 19:39 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage
* 19:37 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 16s)
* 19:37 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2017.codfw.wmnet with reason: host reimage
* 19:34 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1026.eqiad.wmnet with reason: host reimage
* 19:33 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 12s)
* 19:33 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 19:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1026.eqiad.wmnet with OS bookworm
* 19:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2017
* 19:16 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2017
* 19:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2017.codfw.wmnet with OS bookworm
* 18:30 bking@dns1004: END - running authdns-update
* 18:28 bking@dns1004: START - running authdns-update
* 18:16 kamila@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1264.eqiad.wmnet
* 18:16 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1264.eqiad.wmnet
* 18:16 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1264.eqiad.wmnet
* 17:49 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 17:46 dzahn@dns1006: END - running authdns-update
* 17:44 dzahn@dns1006: START - running authdns-update
* 17:44 dzahn@dns1006: END - running authdns-update
* 17:42 dzahn@dns1006: START - running authdns-update
* 17:28 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 17:21 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264
* 17:01 kamila@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264
* 17:01 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 17:01 kamila@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet
* 17:01 kamila@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet
* 17:01 kamila@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet
* 16:42 reedy@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] (duration: 10m 29s)
* 16:34 reedy@deploy2003: reedy, hartman: Continuing with deployment
* 16:33 reedy@deploy2003: reedy, hartman: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:31 reedy@deploy2003: Started scap sync-world: Backport for [[gerrit:1311872{{!}}generatePngAndMidi.sh: Remove SCORE_SAFE (T432418)]]
* 16:26 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 16:10 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in2001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 16:07 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-in1001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 16:05 kamila@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 16:01 kamila@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1264.eqiad.wmnet with reason: host reimage
* 16:00 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out2001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 15:41 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 15:41 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 15:35 jhathaway@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on mx-out1001.wikimedia.org with reason: [[phab:T431659|T431659]]
* 15:14 cgoubert@cumin2003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1339.eqiad.wmnet
* 15:13 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1339.eqiad.wmnet
* 15:13 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1339.eqiad.wmnet
* 14:51 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:50 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:49 kamila@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:49 kamila@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:33 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 14:27 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1339.eqiad.wmnet with reason: host reimage
* 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1339
* 14:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1339
* 14:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1339
* 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1339.eqiad.wmnet 156.32.64.10.in-addr.arpa 6.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003"
* 14:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1339 - cgoubert@cumin2003"
* 14:09 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1339
* 14:06 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1339.eqiad.wmnet with OS trixie
* 14:03 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1339.eqiad.wmnet
* 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1339.eqiad.wmnet
* 14:02 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1339.eqiad.wmnet
* 13:45 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1013.eqiad.wmnet
* 13:39 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1013.eqiad.wmnet
* 13:27 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99)
* 13:24 blake@dns1004: END - running authdns-update
* 13:22 blake@dns1004: START - running authdns-update
* 13:20 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 13:11 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2014.codfw.wmnet
* 13:06 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2014.codfw.wmnet
* 13:06 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb2012.codfw.wmnet
* 13:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 13:03 cwilliams@cumin1003: START - Cookbook sre.mysql.update-replication
* 13:01 fnegri@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for 15 hosts
* 13:01 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb2012.codfw.wmnet
* 13:01 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1016.eqiad.wmnet
* 13:00 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for 15 hosts
* 12:55 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1016.eqiad.wmnet
* 12:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rdb1014.eqiad.wmnet
* 12:49 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host rdb1014.eqiad.wmnet
* 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:32 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:31 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1338.eqiad.wmnet
* 12:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1338.eqiad.wmnet
* 12:18 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1338.eqiad.wmnet
* 12:17 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:15 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:14 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:13 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1338.eqiad.wmnet with OS trixie
* 12:01 klausman@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 11:59 klausman@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 11:56 klausman@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 11:54 klausman@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 11:53 klausman@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 11:51 klausman@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 11:42 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage
* 11:38 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1338.eqiad.wmnet with reason: host reimage
* 11:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet
* 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1338
* 11:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1338
* 11:25 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1338
* 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:25 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1338.eqiad.wmnet 155.32.64.10.in-addr.arpa 5.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:25 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003"
* 11:25 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1338 - cgoubert@cumin2003"
* 11:23 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet
* 11:20 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1338
* 11:20 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1338.eqiad.wmnet with OS trixie
* 11:19 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1338.eqiad.wmnet
* 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1338.eqiad.wmnet
* 11:19 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1338.eqiad.wmnet
* 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1337.eqiad.wmnet
* 11:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1337.eqiad.wmnet
* 11:17 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1337.eqiad.wmnet
* 11:02 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1337.eqiad.wmnet with OS trixie
* 10:51 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet
* 10:50 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:43 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage
* 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 10:39 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:39 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1337.eqiad.wmnet with reason: host reimage
* 10:39 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1001-1003].eqiad.wmnet
* 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:34 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:30 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:28 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1001-1003].eqiad.wmnet
* 10:27 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1337
* 10:27 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1337
* 10:26 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1337
* 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:26 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1337.eqiad.wmnet 154.32.64.10.in-addr.arpa 4.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:26 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003"
* 10:26 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1337 - cgoubert@cumin2003"
* 10:21 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 10:18 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1337
* 10:17 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1337.eqiad.wmnet with OS trixie
* 10:17 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1337.eqiad.wmnet
* 10:16 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet
* 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1337.eqiad.wmnet
* 10:16 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1337.eqiad.wmnet
* 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1336.eqiad.wmnet
* 10:15 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1336.eqiad.wmnet
* 10:15 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1336.eqiad.wmnet
* 10:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet
* 10:10 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db1176.eqiad.wmnet
* 10:09 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db1176.eqiad.wmnet
* 10:05 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 13 hosts (check the cookbook's logs for more details.)
* 10:03 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 13 hosts (check the cookbook's logs for more details.)
* 09:58 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1336.eqiad.wmnet with OS trixie
* 09:47 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for 1 hosts (check the cookbook's logs for more details.)
* 09:47 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for 1 hosts (check the cookbook's logs for more details.)
* 09:45 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1
* 09:40 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host acmechief-test2001.codfw.wmnet,acmechief-test1001.eqiad.wmnet,an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet,db-test[2001-2002].codfw.wmnet,db-test[1001-1003].eqiad.wmn
* 09:39 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage
* 09:33 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1336.eqiad.wmnet with reason: host reimage
* 09:29 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet
* 09:29 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet
* 09:28 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:26 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:21 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1336
* 09:20 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1336
* 09:19 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 09:14 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1336
* 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1336.eqiad.wmnet 152.32.64.10.in-addr.arpa 2.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:14 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003"
* 09:14 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1336 - cgoubert@cumin2003"
* 09:11 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:10 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 09:09 elukey@cumin2002: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1336
* 09:09 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1336.eqiad.wmnet with OS trixie
* 09:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1336.eqiad.wmnet
* 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1336.eqiad.wmnet
* 09:08 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1336.eqiad.wmnet
* 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1335.eqiad.wmnet
* 09:06 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1335.eqiad.wmnet
* 09:06 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1335.eqiad.wmnet
* 09:04 elukey: uploaded spicerack_13.1.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 08:55 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp2001.codfw.wmnet
* 08:54 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host testreduce1002.eqiad.wmnet
* 08:52 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1335.eqiad.wmnet with OS trixie
* 08:51 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp2001.codfw.wmnet
* 08:51 jiji@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wikikube-worker-exp1001.eqiad.wmnet
* 08:50 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host testreduce1002.eqiad.wmnet
* 08:45 jiji@cumin2003: START - Cookbook sre.hosts.reboot-single for host wikikube-worker-exp1001.eqiad.wmnet
* 08:34 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage
* 08:30 cgoubert@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1335.eqiad.wmnet with reason: host reimage
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1335
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1335
* 08:18 cgoubert@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1335
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:18 cgoubert@cumin2003: START - Cookbook sre.dns.wipe-cache wikikube-worker1335.eqiad.wmnet 150.32.64.10.in-addr.arpa 0.5.1.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:18 cgoubert@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003"
* 08:18 cgoubert@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1335 - cgoubert@cumin2003"
* 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 08:14 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges
* 08:13 cgoubert@cumin2003: START - Cookbook sre.dns.netbox
* 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1335
* 08:10 cgoubert@cumin2003: START - Cookbook sre.hosts.reimage for host wikikube-worker1335.eqiad.wmnet with OS trixie
* 08:09 cgoubert@cumin2003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1335.eqiad.wmnet
* 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1335.eqiad.wmnet
* 08:09 cgoubert@cumin2003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1335.eqiad.wmnet
* 08:06 elukey@cumin1003: END (FAIL) - Cookbook sre.puppet.disable-merges (exit_code=99)
* 08:05 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges
* 08:03 elukey@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin1003.eqiad.wmnet
* 07:57 elukey@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin1003.eqiad.wmnet
* 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb1003.eqiad.wmnet
* 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb1003.eqiad.wmnet
* 07:46 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetdb2003.codfw.wmnet
* 07:37 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetdb2003.codfw.wmnet
* 07:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1001.eqiad.wmnet
* 07:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1001.eqiad.wmnet
* 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 07:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2002.codfw.wmnet
* 07:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2002.codfw.wmnet
* 07:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 07:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 07:04 btullis@cumin1003: END (FAIL) - Cookbook sre.hadoop.reboot-workers (exit_code=99) for Hadoop analytics cluster
* 07:04 elukey@cumin1003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 07:04 elukey@cumin1003: START - Cookbook sre.puppet.disable-merges
* 06:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox1003.eqiad.wmnet
* 06:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox1003.eqiad.wmnet
* 02:46 ryankemper@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:46 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:44 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:37 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore categories on wdqs1025 after Bookworm reimage) xfer categories from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling both afterwards
* 02:37 ryankemper@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 49s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 01:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards
* 01:07 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards
* 00:55 urbanecm@deploy2003: helmfile [codfw] DONE helmfile.d/services/linkrecommendation: apply
* 00:54 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/linkrecommendation: apply
* 00:54 urbanecm@deploy2003: helmfile [staging] DONE helmfile.d/services/linkrecommendation: apply
* 00:54 urbanecm@deploy2003: helmfile [codfw] START helmfile.d/services/linkrecommendation: apply
* 00:53 urbanecm@deploy2003: helmfile [staging] START helmfile.d/services/linkrecommendation: apply
* 00:52 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/linkrecommendation: apply
* 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs1025.eqiad.wmnet, repooling source-only afterwards
* 00:23 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], Restore wdqs1027 after Bookworm reimage) xfer scholarly_articles from wdqs2026.codfw.wmnet -> wdqs1027.eqiad.wmnet, repooling both afterwards
* 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1274.eqiad.wmnet
* 00:14 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1274.eqiad.wmnet
* 00:14 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1274.eqiad.wmnet
* 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1027.eqiad.wmnet with OS bookworm
* 00:13 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs1025.eqiad.wmnet with OS bookworm
* 00:04 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1274.eqiad.wmnet with OS trixie
== 2026-07-16 ==
* 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards
* 23:50 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage
* 23:47 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage
* 23:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage
* 23:41 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1025.eqiad.wmnet with reason: host reimage
* 23:39 ryankemper@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs1027.eqiad.wmnet with reason: host reimage
* 23:38 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1274.eqiad.wmnet with reason: host reimage
* 23:23 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1025
* 23:23 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1025
* 23:22 ryankemper@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs1027
* 23:22 ryankemper@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs1027
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1274
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1274
* 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1025.eqiad.wmnet with OS bookworm
* 23:19 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1274
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:19 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1274.eqiad.wmnet 145.48.64.10.in-addr.arpa 5.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:19 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003"
* 23:19 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1274 - swfrench@cumin1003"
* 23:19 ryankemper@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs1027.eqiad.wmnet with OS bookworm
* 23:14 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 23:14 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1274
* 23:13 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1274.eqiad.wmnet with OS trixie
* 23:13 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1274.eqiad.wmnet
* 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1274.eqiad.wmnet
* 23:12 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1274.eqiad.wmnet
* 23:12 ryankemper: [[phab:T430880|T430880]] depooled dnsdisc of wdqs-internal-scholarly-eqiad bc we only have 1 host there
* 23:09 ryankemper@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=wdqs-internal-scholarly,name=eqiad
* 23:08 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1272.eqiad.wmnet
* 23:08 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1272.eqiad.wmnet
* 23:08 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1272.eqiad.wmnet
* 23:01 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430880|T430880]], xfer to freshly reimaged/scap-deployed wdqs2025 after Bookworm reimage) xfer wikidata_main from wdqs2020.codfw.wmnet -> wdqs2025.codfw.wmnet, repooling source-only afterwards
* 22:57 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1272.eqiad.wmnet with OS trixie
* 22:56 Amir1: deleting echo notifications from 2015 in group0
* 22:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wdqs2025.codfw.wmnet with OS bookworm
* 22:35 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage
* 22:32 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]] (duration: 00m 27s)
* 22:32 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c]: [[phab:T430880|T430880]]
* 22:28 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1269.eqiad.wmnet
* 22:28 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1269.eqiad.wmnet
* 22:28 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1269.eqiad.wmnet
* 22:27 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1272.eqiad.wmnet with reason: host reimage
* 22:26 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] (duration: 08m 51s)
* 22:22 ladsgroup@deploy2003: ladsgroup, urbanecm: Continuing with deployment
* 22:19 ladsgroup@deploy2003: ladsgroup, urbanecm: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:17 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1311420{{!}}Upload: Do not throw for failure to save a chunk file (T430986)]], [[gerrit:1311529{{!}}AddImage: Request only standard thumbnail sizes (T428797)]]
* 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage
* 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1272
* 22:06 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1272
* 22:05 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1272
* 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:05 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1272.eqiad.wmnet 127.48.64.10.in-addr.arpa 7.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:05 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003"
* 22:05 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1272 - swfrench@cumin1003"
* 22:03 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wdqs2025.codfw.wmnet with reason: host reimage
* 22:01 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1272
* 22:00 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1272.eqiad.wmnet with OS trixie
* 22:00 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1272.eqiad.wmnet
* 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1272.eqiad.wmnet
* 21:59 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1272.eqiad.wmnet
* 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1271.eqiad.wmnet
* 21:55 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1271.eqiad.wmnet
* 21:55 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1271.eqiad.wmnet
* 21:46 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1271.eqiad.wmnet with OS trixie
* 21:45 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] (duration: 06m 31s)
* 21:40 sbassett@deploy2003: sbassett: Continuing with deployment
* 21:40 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wdqs2025
* 21:40 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wdqs2025
* 21:40 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:38 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311482{{!}}Use non-sampled authentication log channel instead of authevents (T432042)]]
* 21:37 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wdqs2025
* 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:37 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wdqs2025.codfw.wmnet 220.48.192.10.in-addr.arpa 0.2.2.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:37 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:37 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003"
* 21:37 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wdqs2025 - bking@cumin2003"
* 21:30 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] (duration: 08m 19s)
* 21:26 sbassett@deploy2003: sbassett: Continuing with deployment
* 21:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1269.eqiad.wmnet with OS trixie
* 21:24 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage
* 21:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:22 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:22 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1311481{{!}}Cleanup: remove reauth indicator from log message (T432042)]]
* 21:20 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wdqs2025
* 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wdqs2025.codfw.wmnet with OS bookworm
* 21:17 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1271.eqiad.wmnet with reason: host reimage
* 21:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage
* 21:00 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1269.eqiad.wmnet with reason: host reimage
* 20:56 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1271
* 20:55 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1271
* 20:54 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1271
* 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:54 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1271.eqiad.wmnet 126.48.64.10.in-addr.arpa 6.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:54 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003"
* 20:54 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1271 - swfrench@cumin1003"
* 20:51 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:51 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:51 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:50 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:49 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 20:49 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1268.eqiad.wmnet
* 20:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1268.eqiad.wmnet
* 20:49 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1268.eqiad.wmnet
* 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1271
* 20:48 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1271.eqiad.wmnet with OS trixie
* 20:47 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1271.eqiad.wmnet
* 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1271.eqiad.wmnet
* 20:46 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1271.eqiad.wmnet
* 20:41 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] (duration: 07m 34s)
* 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1269
* 20:39 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1269
* 20:36 aude@deploy2003: aude: Continuing with deployment
* 20:35 aude@deploy2003: aude: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:33 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1311472{{!}}Preserve menus after toolbox (e.g. print/export) in page tools (T432316)]]
* 20:26 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on db2207.codfw.wmnet with reason: Host down
* 20:22 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-video: apply
* 20:21 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-video: apply
* 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-timeline: apply
* 20:20 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-timeline: apply
* 20:20 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 20:19 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 20:19 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-media: apply
* 20:18 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-media: apply
* 20:18 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox-constraints: apply
* 20:17 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox-constraints: apply
* 20:17 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/shellbox: apply
* 20:16 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/shellbox: apply
* 20:13 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1269
* 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:13 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1269.eqiad.wmnet 80.32.64.10.in-addr.arpa 0.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 20:13 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003"
* 20:13 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1269 - kamila@cumin1003"
* 20:09 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons.
* 20:07 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 20:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 20:03 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-codfw cluster: Roll restart of jvm daemons.
* 20:03 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2207 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94893 and previous config saved to /var/cache/conftool/dbconfig/20260716-200257-marostegui.json
* 20:01 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2204 to s2 primary [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94892 and previous config saved to /var/cache/conftool/dbconfig/20260716-200157-marostegui.json
* 20:00 marostegui: Starting emergency s2 codfw failover from db2207 to db2204 - [[phab:T432396|T432396]]
* 19:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1035.eqiad.wmnet
* 19:56 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2204 with weight 0 [[phab:T432396|T432396]]', diff saved to https://phabricator.wikimedia.org/P94891 and previous config saved to /var/cache/conftool/dbconfig/20260716-195628-marostegui.json
* 19:55 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: Primary switchover s2 [[phab:T432396|T432396]]
* 19:54 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1035.eqiad.wmnet
* 19:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1034.eqiad.wmnet
* 19:48 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1034.eqiad.wmnet
* 19:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1033.eqiad.wmnet
* 19:43 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-video: apply
* 19:43 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1033.eqiad.wmnet
* 19:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1032.eqiad.wmnet
* 19:42 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-video: apply
* 19:42 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-timeline: apply
* 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-timeline: apply
* 19:41 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:41 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:40 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-media: apply
* 19:40 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-media: apply
* 19:39 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox-constraints: apply
* 19:36 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox-constraints: apply
* 19:36 swfrench@deploy2003: helmfile [codfw] DONE helmfile.d/services/shellbox: apply
* 19:35 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1032.eqiad.wmnet
* 19:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1031.eqiad.wmnet
* 19:35 swfrench@deploy2003: helmfile [codfw] START helmfile.d/services/shellbox: apply
* 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-video: apply
* 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-video: apply
* 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-timeline: apply
* 19:33 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-timeline: apply
* 19:33 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-syntaxhighlight: apply
* 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-media: apply
* 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-media: apply
* 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox-constraints: apply
* 19:32 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox-constraints: apply
* 19:32 swfrench@deploy2003: helmfile [staging] DONE helmfile.d/services/shellbox: apply
* 19:31 swfrench@deploy2003: helmfile [staging] START helmfile.d/services/shellbox: apply
* 19:27 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1031.eqiad.wmnet
* 19:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1030.eqiad.wmnet
* 19:23 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards
* 19:19 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1030.eqiad.wmnet
* 19:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host wdqs1029.eqiad.wmnet
* 19:17 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 19:12 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host wdqs1029.eqiad.wmnet
* 19:06 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1269
* 19:05 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1269.eqiad.wmnet with OS trixie
* 19:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1269.eqiad.wmnet
* 19:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1269.eqiad.wmnet
* 19:03 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1269.eqiad.wmnet
* 18:55 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 02m 41s)
* 18:53 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]]
* 18:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1268.eqiad.wmnet with OS trixie
* 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 18:16 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1267.eqiad.wmnet
* 18:16 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1267.eqiad.wmnet
* 18:16 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1267.eqiad.wmnet
* 18:09 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage
* 18:08 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1001.eqiad.wmnet, repooling source-only afterwards
* 18:06 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] (duration: 07m 34s)
* 18:06 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 23s)
* 18:06 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 18:06 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1268.eqiad.wmnet with reason: host reimage
* 18:03 bd808@deploy2003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 18:02 bd808@deploy2003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 18:02 swfrench@deploy2003: jiji, swfrench: Continuing with deployment
* 18:02 bd808@deploy2003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 18:02 bd808@deploy2003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 18:01 bd808@deploy2003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 18:01 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:01 bd808@deploy2003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 17:59 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311468{{!}}ProductionServices: repool poolcounter2006 after reboot (T431705)]]
* 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1268
* 17:45 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1268
* 17:44 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1267.eqiad.wmnet with OS trixie
* 17:43 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2006.codfw.wmnet
* 17:39 swfrench@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2006.codfw.wmnet
* 17:35 swfrench@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] (duration: 07m 27s)
* 17:34 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1268
* 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:34 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1268.eqiad.wmnet 78.32.64.10.in-addr.arpa 8.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:34 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003"
* 17:34 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1268 - kamila@cumin1003"
* 17:31 swfrench@deploy2003: jiji, swfrench: Continuing with deployment
* 17:29 swfrench@deploy2003: jiji, swfrench: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:28 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 17:28 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1268
* 17:28 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1268.eqiad.wmnet with OS trixie
* 17:27 swfrench@deploy2003: Started scap sync-world: Backport for [[gerrit:1311467{{!}}ProductionServices: reboot poolcounter2006 (#4/4) (T431705)]]
* 17:23 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage
* 17:18 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1267.eqiad.wmnet with reason: host reimage
* 17:18 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1268.eqiad.wmnet
* 17:17 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1268.eqiad.wmnet
* 17:17 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1268.eqiad.wmnet
* 17:12 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter2005.codfw.wmnet
* 17:11 swfrench@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1270.eqiad.wmnet
* 17:11 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1270.eqiad.wmnet
* 17:11 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1270.eqiad.wmnet
* 17:09 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter2005.codfw.wmnet
* 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 17:08 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003"
* 17:08 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: update reverse dns for moved arelion cct cr2-eqiad - cmooney@cumin1003"
* 17:08 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] (duration: 07m 34s)
* 17:04 jiji@deploy2003: jiji: Continuing with deployment
* 17:03 jiji@deploy2003: jiji: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:01 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 17:00 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311466{{!}}ProductionServices: reboot poolcounter2005 (#3/4) (T431705)]]
* 17:00 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2185.codfw.wmnet with OS trixie
* 16:59 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 16:58 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1270.eqiad.wmnet with OS trixie
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1267
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1267
* 16:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1267
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:57 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1267.eqiad.wmnet 77.32.64.10.in-addr.arpa 7.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:57 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003"
* 16:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1267 - kamila@cumin1003"
* 16:56 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqsin
* 16:56 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5032.eqsin.wmnet
* 16:52 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_esams
* 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3073.esams.wmnet
* 16:50 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_esams
* 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3081.esams.wmnet
* 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1266.eqiad.wmnet
* 16:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1266.eqiad.wmnet
* 16:45 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1266.eqiad.wmnet
* 16:43 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2185.codfw.wmnet with reason: host reimage
* 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqiad
* 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1114.eqiad.wmnet
* 16:41 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_eqiad
* 16:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1115.eqiad.wmnet
* 16:39 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 16:39 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1267
* 16:39 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2185.codfw.wmnet with reason: host reimage
* 16:38 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1267.eqiad.wmnet with OS trixie
* 16:38 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1267.eqiad.wmnet
* 16:38 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage
* 16:37 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1267.eqiad.wmnet
* 16:37 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1267.eqiad.wmnet
* 16:31 swfrench@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1270.eqiad.wmnet with reason: host reimage
* 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1264.eqiad.wmnet
* 16:24 kamila@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 16:24 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 16:23 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 16:21 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: switch maintenance completed codfw rack b6
* 16:20 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2185.codfw.wmnet with OS trixie
* 16:19 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 16:16 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1007.eqiad.wmnet
* 16:15 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_eqsin
* 16:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5024.eqsin.wmnet
* 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5031.eqsin.wmnet
* 16:13 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3072.esams.wmnet
* 16:12 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1007.eqiad.wmnet
* 16:11 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] (duration: 09m 47s)
* 16:10 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1266.eqiad.wmnet with OS trixie
* 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1270
* 16:10 swfrench@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1270
* 16:09 swfrench@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1270
* 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:09 swfrench@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1270.eqiad.wmnet 125.48.64.10.in-addr.arpa 5.2.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:09 swfrench@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003"
* 16:09 swfrench@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1270 - swfrench@cumin1003"
* 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3080.esams.wmnet
* 16:07 jiji@deploy2003: jiji: Continuing with deployment
* 16:06 jiji@deploy2003: jiji: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:04 swfrench@cumin1003: START - Cookbook sre.dns.netbox
* 16:04 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1265.eqiad.wmnet
* 16:03 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1265.eqiad.wmnet
* 16:03 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1265.eqiad.wmnet
* 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1270
* 16:03 swfrench@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1270.eqiad.wmnet with OS trixie
* 16:02 swfrench@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1270.eqiad.wmnet
* 16:02 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311465{{!}}ProductionServices: reboot poolcounter1007 (#2/4) (T431705)]]
* 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1270.eqiad.wmnet
* 16:01 swfrench@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1270.eqiad.wmnet
* 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1113.eqiad.wmnet
* 16:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1112.eqiad.wmnet
* 15:49 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage
* 15:47 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host poolcounter1006.eqiad.wmnet
* 15:45 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1265.eqiad.wmnet with OS trixie
* 15:44 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1266.eqiad.wmnet with reason: host reimage
* 15:43 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host poolcounter1006.eqiad.wmnet
* 15:42 jiji@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] (duration: 09m 46s)
* 15:37 jiji@deploy2003: jiji: Continuing with deployment
* 15:36 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2162: switch maintenance completed codfw rack b6
* 15:36 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2161: switch maintenance completed codfw rack b6
* 15:34 jiji@deploy2003: jiji: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 15:32 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5023.eqsin.wmnet
* 15:32 jiji@deploy2003: Started scap sync-world: Backport for [[gerrit:1311464{{!}}ProductionServices: reboot poolcounter1006 (#1/4) (T431705)]]
* 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5030.eqsin.wmnet
* 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3071.esams.wmnet
* 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3079.esams.wmnet
* 15:25 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage
* 15:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1266
* 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1266
* 15:21 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1110.eqiad.wmnet
* 15:20 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1111.eqiad.wmnet
* 15:16 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1265.eqiad.wmnet with reason: host reimage
* 15:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1001.eqiad.wmnet with OS bookworm
* 15:15 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1266
* 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:15 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1266.eqiad.wmnet 76.32.64.10.in-addr.arpa 6.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:15 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003"
* 15:15 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1266 - kamila@cumin1003"
* 15:07 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 15:04 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1266
* 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1264
* 15:04 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1264
* 15:04 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1266.eqiad.wmnet with OS trixie
* 15:03 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1264
* 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:03 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1264.eqiad.wmnet 74.32.64.10.in-addr.arpa 4.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:03 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003"
* 15:03 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1264 - kamila@cumin1003"
* 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-eqiad
* 15:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:01 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp1001.eqiad.wmnet
* 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1376-1384].eqiad.wmnet
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1376-1384].eqiad.wmnet
* 14:59 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet
* 14:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1266.eqiad.wmnet
* 14:58 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1266.eqiad.wmnet
* 14:58 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1266.eqiad.wmnet
* 14:58 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 14:57 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1264
* 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1265
* 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1265
* 14:57 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1310596{{!}}Set $wgMathInternalRestbaseURL explicitly (T349582)]]
* 14:57 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1265
* 14:57 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:56 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1265.eqiad.wmnet 75.32.64.10.in-addr.arpa 5.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:56 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003"
* 14:56 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1265 - kamila@cumin1003"
* 14:53 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet
* 14:53 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1376-1384].eqiad.wmnet
* 14:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage
* 14:51 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet
* 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet
* 14:50 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:50 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2161: switch maintenance completed codfw rack b6
* 14:50 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5021.eqsin.wmnet
* 14:50 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1265
* 14:49 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1264.eqiad.wmnet with OS trixie
* 14:49 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1265.eqiad.wmnet with OS trixie
* 14:49 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5029.eqsin.wmnet
* 14:49 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1265.eqiad.wmnet
* 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3070.esams.wmnet
* 14:48 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1264.eqiad.wmnet
* 14:48 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1001.eqiad.wmnet with reason: host reimage
* 14:48 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1376-1384].eqiad.wmnet
* 14:48 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1265.eqiad.wmnet
* 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1265.eqiad.wmnet
* 14:47 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1264.eqiad.wmnet
* 14:47 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1264.eqiad.wmnet
* 14:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:47 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3078.esams.wmnet
* 14:44 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1263.eqiad.wmnet
* 14:44 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1263.eqiad.wmnet
* 14:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1263.eqiad.wmnet
* 14:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1108.eqiad.wmnet
* 14:40 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1109.eqiad.wmnet
* 14:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1367-1375].eqiad.wmnet
* 14:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:34 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2006.codfw.wmnet
* 14:34 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons.
* 14:32 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2002.codfw.wmnet
* 14:31 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1002.eqiad.wmnet
* 14:29 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2006.codfw.wmnet
* 14:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:27 kamila@deploy2003: Finished scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 02m 57s)
* 14:27 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-flink-eqiad cluster: Roll restart of jvm daemons.
* 14:26 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2002.codfw.wmnet
* 14:26 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf2001.codfw.wmnet
* 14:25 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1002.eqiad.wmnet
* 14:25 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-wf1001.eqiad.wmnet
* 14:25 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]]
* 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1358-1366].eqiad.wmnet
* 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:21 kamila@deploy2003: sync-world aborted: Test deployment to check rsync is working - [[phab:T432108|T432108]] (duration: 00m 36s)
* 14:21 topranks: reboot lsw1-b6-codfw to upgrade JunOS [[phab:T430922|T430922]]
* 14:21 kamila@deploy2003: Started scap sync-world: Test deployment to check rsync is working - [[phab:T432108|T432108]]
* 14:21 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1001.eqiad.wmnet with OS bookworm
* 14:20 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b6-codfw,lsw1-b6-codfw IPv6,lsw1-b6-codfw.mgmt,ssw1-a[1,8]-codfw with reason: lsw1-b6-codfw JunOS upgrade
* 14:20 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf2001.codfw.wmnet
* 14:19 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-wf1001.eqiad.wmnet
* 14:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 26 hosts with reason: lsw1-b6-codfw JunOS upgrade
* 14:14 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2006.codfw.wmnet
* 14:13 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2006.codfw.wmnet
* 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc2022: switch maintenance codfw rack b6
* 14:13 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool pc2022: switch maintenance codfw rack b6
* 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2251: switch maintenance codfw rack b6
* 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2251: switch maintenance codfw rack b6
* 14:12 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2162: switch maintenance codfw rack b6
* 14:12 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1263.eqiad.wmnet with OS trixie
* 14:12 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2162: switch maintenance codfw rack b6
* 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2161: switch maintenance codfw rack b6
* 14:11 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2161: switch maintenance codfw rack b6
* 14:08 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2093-2094,2102-2106,2279-2283].codfw.wmnet
* 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5020.eqsin.wmnet
* 14:07 btullis@cumin1003: START - Cookbook sre.hadoop.reboot-workers for Hadoop analytics cluster
* 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3069.esams.wmnet
* 14:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1348-1357].eqiad.wmnet
* 14:07 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5028.eqsin.wmnet
* 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1338-1347].eqiad.wmnet
* 14:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1338-1347].eqiad.wmnet
* 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3077.esams.wmnet
* 14:02 topranks: beginning depools for lsw1-b6-codfw maintenance [[phab:T430922|T430922]]
* 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1106.eqiad.wmnet
* 14:00 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1002.eqiad.wmnet
* 13:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1338-1347].eqiad.wmnet
* 13:59 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1107.eqiad.wmnet
* 13:56 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-codfw
* 13:55 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 13:54 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1002.eqiad.wmnet
* 13:54 jiji@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host mc-misc1001.eqiad.wmnet
* 13:54 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 13:50 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage
* 13:50 sfaci@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1338-1347].eqiad.wmnet
* 13:49 jiji@cumin1003: START - Cookbook sre.hosts.reboot-single for host mc-misc1001.eqiad.wmnet
* 13:49 sfaci@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 13:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:49 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:45 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1263.eqiad.wmnet with reason: host reimage
* 13:40 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:40 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-eqiad
* 13:35 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1328-1337].eqiad.wmnet
* 13:34 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:33 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox)
* 13:33 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns7001.wikimedia.org
* 13:27 sukhe@dns1004: END - running authdns-update
* 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5019.eqsin.wmnet
* 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3076.esams.wmnet
* 13:25 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3068.esams.wmnet
* 13:25 sukhe@dns1004: START - running authdns-update
* 13:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5027.eqsin.wmnet
* 13:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1263
* 13:24 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1263
* 13:23 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1263
* 13:23 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:23 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1104.eqiad.wmnet
* 13:21 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:21 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:20 kamila@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:20 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1263.eqiad.wmnet 73.32.64.10.in-addr.arpa 3.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:20 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003"
* 13:20 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1263 - kamila@cumin1003"
* 13:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1105.eqiad.wmnet
* 13:19 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1318-1327].eqiad.wmnet
* 13:19 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:18 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns7001.wikimedia.org
* 13:16 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*
* 13:14 cdobbins@dns1004: END - running authdns-update
* 13:13 sbisson@deploy2003: Finished scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] (duration: 08m 03s)
* 13:13 cdobbins@dns1004: START - running authdns-update
* 13:12 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 13:12 cdobbins@cumin2003: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update
* 13:12 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1263
* 13:11 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1263.eqiad.wmnet with OS trixie
* 13:11 cdobbins@cumin2003: conftool action : set/pooled=no; selector: name=dns7002.*
* 13:11 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1263.eqiad.wmnet
* 13:10 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1263.eqiad.wmnet
* 13:10 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1263.eqiad.wmnet
* 13:09 sbisson@deploy2003: sbisson: Continuing with deployment
* 13:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:07 sbisson@deploy2003: sbisson: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:05 sbisson@deploy2003: Started scap sync-world: Backport for [[gerrit:1311054{{!}}Enable Article Guidance on Polish Wikipedia (T432137)]]
* 13:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6002.wikimedia.org
* 13:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1308-1317].eqiad.wmnet
* 13:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1298-1307].eqiad.wmnet
* 13:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1298-1307].eqiad.wmnet
* 12:59 kamila@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1262.eqiad.wmnet
* 12:59 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1262.eqiad.wmnet
* 12:59 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1262.eqiad.wmnet
* 12:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1298-1307].eqiad.wmnet
* 12:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6002.wikimedia.org
* 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1298-1307].eqiad.wmnet
* 12:46 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:46 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3075.esams.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3067.esams.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5018.eqsin.wmnet
* 12:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5026.eqsin.wmnet
* 12:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1102.eqiad.wmnet
* 12:39 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1103.eqiad.wmnet
* 12:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns6001.wikimedia.org
* 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1287-1289,1291-1297].eqiad.wmnet
* 12:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:18 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns6001.wikimedia.org
* 12:14 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1277-1286].eqiad.wmnet
* 12:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1267-1276].eqiad.wmnet
* 12:13 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1267-1276].eqiad.wmnet
* 12:04 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1267-1276].eqiad.wmnet
* 12:03 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5004.wikimedia.org
* 12:02 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1100.eqiad.wmnet
* 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3066.esams.wmnet
* 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp3074.esams.wmnet
* 12:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:01 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-codfw
* 12:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5017.eqsin.wmnet
* 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp5025.eqsin.wmnet
* 12:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp1101.eqiad.wmnet
* 11:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1267-1276].eqiad.wmnet
* 11:58 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:58 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:54 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5004.wikimedia.org
* 11:54 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and (A:eqsin or A:drmrs or A:magru) and not (P<nowiki>{</nowiki>dns5003*<nowiki>}</nowiki> or P<nowiki>{</nowiki>dns7002*<nowiki>}</nowiki>) and (A:dnsbox)
* 11:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:53 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:53 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-eqiad
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqiad
* 11:50 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:50 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqiad
* 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_esams
* 11:49 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_esams
* 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_eqsin
* 11:48 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_eqsin
* 11:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:44 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-eqiad
* 11:43 jiji@cumin1003: END (PASS) - Cookbook sre.memcached.roll-reboot-restart (exit_code=0) rolling reboot on A:memcached-gutter-codfw
* 11:42 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1256-1261,1263-1266].eqiad.wmnet
* 11:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:24 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-eqiad
* 11:23 jiji@cumin1003: START - Cookbook sre.memcached.roll-reboot-restart rolling reboot on A:memcached-gutter-codfw
* 11:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1246-1255].eqiad.wmnet
* 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2066.codfw.wmnet with OS trixie
* 11:05 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1161-1163,1165,1240-1245].eqiad.wmnet
* 11:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1151-1160].eqiad.wmnet
* 11:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1151-1160].eqiad.wmnet
* 10:59 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1068.eqiad.wmnet with OS trixie
* 10:55 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.major-upgrade (exit_code=99)
* 10:55 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 10:54 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1151-1160].eqiad.wmnet
* 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage
* 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1151-1160].eqiad.wmnet
* 10:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:47 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2066.codfw.wmnet with reason: host reimage
* 10:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:36 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1141-1150].eqiad.wmnet
* 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:22 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2066.codfw.wmnet with OS trixie
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:21 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1131-1140].eqiad.wmnet
* 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:07 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2065.codfw.wmnet with OS trixie
* 10:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:06 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 10:06 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 10:04 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage
* 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:03 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1121-1130].eqiad.wmnet
* 10:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 10:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 09:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1068.eqiad.wmnet with reason: host reimage
* 09:57 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 09:57 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 09:52 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 09:47 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:46 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage
* 09:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2065.codfw.wmnet with reason: host reimage
* 09:40 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1094-1095,1113-1120].eqiad.wmnet
* 09:39 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:39 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:39 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:37 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox2003.codfw.wmnet
* 09:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox2003.codfw.wmnet
* 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:25 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1077-1081,1084-1087,1093].eqiad.wmnet
* 09:24 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:24 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:21 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2065.codfw.wmnet with OS trixie
* 09:13 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1068.eqiad.wmnet with OS trixie
* 09:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2064.codfw.wmnet with OS trixie
* 09:08 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1067-1076].eqiad.wmnet
* 09:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 09:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 09:07 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2162: Repooling after switchover
* 09:01 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1067.eqiad.wmnet with OS trixie
* 09:00 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:59 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 08:57 tappof: bump space for prometheus k8s-dse in eqiad
* 08:56 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping2004.codfw.wmnet
* 08:52 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping2004.codfw.wmnet
* 08:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1051-1057,1064-1066].eqiad.wmnet
* 08:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ping1004.eqiad.wmnet
* 08:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:51 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:49 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host ping1004.eqiad.wmnet
* 08:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2064.codfw.wmnet with reason: host reimage
* 08:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage
* 08:34 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1041-1050].eqiad.wmnet
* 08:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:33 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:33 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1067.eqiad.wmnet with reason: host reimage
* 08:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:21 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover
* 08:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2064.codfw.wmnet with OS trixie
* 08:16 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1067.eqiad.wmnet with OS trixie
* 08:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[1006-1007,1015-1016,1021,1034-1035,1038-1040].eqiad.wmnet
* 08:15 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-eqiad
* 08:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie
* 08:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1066.eqiad.wmnet with OS trixie
* 08:02 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2162: Repooling after switchover
* 07:56 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2162: Repooling after switchover
* 07:55 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2162 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94870 and previous config saved to /var/cache/conftool/dbconfig/20260716-075530-cwilliams.json
* 07:53 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2241 to x3 primary [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94869 and previous config saved to /var/cache/conftool/dbconfig/20260716-075314-cwilliams.json
* 07:52 cezmunsta: Starting x3 codfw failover from db2162 to db2241 - [[phab:T430925|T430925]]
* 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:50 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 07:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2241 with weight 0 [[phab:T430925|T430925]]', diff saved to https://phabricator.wikimedia.org/P94868 and previous config saved to /var/cache/conftool/dbconfig/20260716-074507-cwilliams.json
* 07:44 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 18 hosts with reason: Primary switchover x3 [[phab:T430925|T430925]]
* 07:43 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage
* 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-eqiad
* 07:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1028.eqiad.wmnet
* 07:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1028.eqiad.wmnet
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 07:38 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1066.eqiad.wmnet with reason: host reimage
* 07:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1028.eqiad.wmnet
* 07:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1028.eqiad.wmnet
* 07:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1027.eqiad.wmnet
* 07:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1027.eqiad.wmnet
* 07:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1027.eqiad.wmnet
* 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1027.eqiad.wmnet
* 07:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1026.eqiad.wmnet
* 07:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1026.eqiad.wmnet
* 07:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast2003.wikimedia.org
* 07:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1026.eqiad.wmnet
* 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1066.eqiad.wmnet with OS trixie
* 07:19 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast2003.wikimedia.org
* 07:18 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie
* 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1026.eqiad.wmnet
* 06:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1025.eqiad.wmnet
* 06:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1025.eqiad.wmnet
* 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 06:47 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 06:44 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1025.eqiad.wmnet
* 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1025.eqiad.wmnet
* 06:14 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1024.eqiad.wmnet
* 06:14 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1024.eqiad.wmnet
* 06:07 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1024.eqiad.wmnet
* 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1024.eqiad.wmnet
* 05:37 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1023.eqiad.wmnet
* 05:37 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1023.eqiad.wmnet
* 05:26 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1023.eqiad.wmnet
* 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1023.eqiad.wmnet
* 04:56 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1022.eqiad.wmnet
* 04:56 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1022.eqiad.wmnet
* 04:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1022.eqiad.wmnet
* 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1022.eqiad.wmnet
* 04:19 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1021.eqiad.wmnet
* 04:19 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1021.eqiad.wmnet
* 04:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1021.eqiad.wmnet
* 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1021.eqiad.wmnet
* 03:38 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1020.eqiad.wmnet
* 03:38 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1020.eqiad.wmnet
* 03:20 btullis@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) depool for host dse-k8s-worker1020.eqiad.wmnet
* 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1020.eqiad.wmnet
* 03:18 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1019.eqiad.wmnet
* 03:18 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1019.eqiad.wmnet
* 03:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1019.eqiad.wmnet
* 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1019.eqiad.wmnet
* 02:41 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1018.eqiad.wmnet
* 02:41 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1018.eqiad.wmnet
* 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards
* 02:36 ryankemper@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards
* 02:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1018.eqiad.wmnet
* 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1018.eqiad.wmnet
* 02:30 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1014.eqiad.wmnet
* 02:30 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1014.eqiad.wmnet
* 02:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1014.eqiad.wmnet
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1014.eqiad.wmnet
* 01:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1013.eqiad.wmnet
* 01:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1013.eqiad.wmnet
* 01:47 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1013.eqiad.wmnet
* 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2003.codfw.wmnet -> wcqs2001.codfw.wmnet, repooling both afterwards
* 01:20 ryankemper@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling both afterwards
* 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1013.eqiad.wmnet
* 01:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1012.eqiad.wmnet
* 01:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1012.eqiad.wmnet
* 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1012.eqiad.wmnet
* 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1012.eqiad.wmnet
* 01:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1011.eqiad.wmnet
* 01:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1011.eqiad.wmnet
* 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage (duration: 00m 23s)
* 01:08 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] scap deploy post bookworm reimage
* 01:08 ryankemper@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage (duration: 00m 46s)
* 01:07 ryankemper@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): scap deploy post bookworm reimage
* 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1011.eqiad.wmnet
* 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1011.eqiad.wmnet
* 01:04 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1010.eqiad.wmnet
* 01:04 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1010.eqiad.wmnet
* 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1010.eqiad.wmnet
* 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1010.eqiad.wmnet
* 00:57 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1009.eqiad.wmnet
* 00:57 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1009.eqiad.wmnet
* 00:50 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1009.eqiad.wmnet
* 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1009.eqiad.wmnet
* 00:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1008.eqiad.wmnet
* 00:20 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1008.eqiad.wmnet
* 00:13 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1008.eqiad.wmnet
== 2026-07-15 ==
* 23:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2001.codfw.wmnet with OS bookworm
* 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1008.eqiad.wmnet
* 23:43 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1007.eqiad.wmnet
* 23:43 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1007.eqiad.wmnet
* 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1007.eqiad.wmnet
* 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1007.eqiad.wmnet
* 23:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1006.eqiad.wmnet
* 23:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1006.eqiad.wmnet
* 23:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1006.eqiad.wmnet
* 23:29 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1006.eqiad.wmnet
* 23:28 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1005.eqiad.wmnet
* 23:28 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1005.eqiad.wmnet
* 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1002.eqiad.wmnet with OS bookworm
* 23:21 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1005.eqiad.wmnet
* 23:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage
* 23:15 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 23:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2001.codfw.wmnet with reason: host reimage
* 23:04 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 23:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 22:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1005.eqiad.wmnet
* 22:51 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1004.eqiad.wmnet
* 22:51 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1004.eqiad.wmnet
* 22:45 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1004.eqiad.wmnet
* 22:44 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 22:44 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS trixie
* 22:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 22:34 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 22:16 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 7 days, 0:00:00 on datahubsearch[1002-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages
* 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1004.eqiad.wmnet
* 22:15 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1003.eqiad.wmnet
* 22:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1003.eqiad.wmnet
* 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1003.eqiad.wmnet
* 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1003.eqiad.wmnet
* 22:08 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1002.eqiad.wmnet
* 22:08 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1002.eqiad.wmnet
* 22:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host datahubsearch1001.eqiad.wmnet with OS bookworm
* 22:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 22:02 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm
* 22:01 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on datahubsearch[1001-1003].eqiad.wmnet with reason: Using datahubsearch1001 to test bookworm reimages
* 22:01 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1002.eqiad.wmnet
* 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1002.eqiad.wmnet
* 22:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker1001.eqiad.wmnet
* 22:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker1001.eqiad.wmnet
* 21:53 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker1001.eqiad.wmnet
* 21:52 swfrench@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 21:50 swfrench@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 21:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie
* 21:50 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:38 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:30 bking@cumin2003: END (FAIL) - Cookbook sre.hardware.upgrade-firmware (exit_code=1) upgrade firmware for hosts ['wcqs1002']
* 21:30 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:29 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 21:29 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:29 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:29 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 21:28 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm
* 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker1001.eqiad.wmnet
* 21:23 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:23 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 21:20 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:18 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf11u2 into component/php83 for bullseye-wikimedia
* 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test1001.eqiad.wmnet
* 21:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:17 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:16 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs1002']
* 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1003.eqiad.wmnet
* 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:11 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:11 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons.
* 21:08 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs1002']
* 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1002.eqiad.wmnet
* 21:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 21:05 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs1001.eqiad.wmnet
* 21:04 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-druid-public cluster: Roll restart of jvm daemons.
* 21:02 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 21:01 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 21:01 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 21:00 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 20:59 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 20:59 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs1001.eqiad.wmnet
* 20:59 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-eqiad
* 20:55 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 20:55 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 20:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm
* 20:21 jhathaway: puppet is re-enabled, have fun, but not too much fun!
* 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 20:17 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001']
* 20:12 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001']
* 20:11 bking@cumin2003: END (PASS) - Cookbook sre.hardware.upgrade-firmware (exit_code=0) upgrade firmware for hosts ['wcqs2001']
* 20:09 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:08 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:05 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:05 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:04 bking@cumin2003: START - Cookbook sre.hardware.upgrade-firmware upgrade firmware for hosts ['wcqs2001']
* 20:03 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 20:03 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host wcqs2001.codfw.wmnet with OS bookworm
* 20:02 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 20:02 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 20:01 jhathaway: disabling puppet fleet wide to roll out kafka patch
* 19:55 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1010.eqiad.wmnet
* 19:52 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 19:52 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 19:48 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 19:48 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 19:48 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 19:47 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 19:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 19:45 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 19:45 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 19:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1010.eqiad.wmnet
* 19:38 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1262.eqiad.wmnet with OS trixie
* 19:17 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage
* 19:11 kamila@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1262.eqiad.wmnet with reason: host reimage
* 18:59 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie
* 18:54 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 18:53 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1262
* 18:52 kamila@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1262
* 18:51 kamila@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1262
* 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:51 kamila@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1262.eqiad.wmnet 72.32.64.10.in-addr.arpa 2.7.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:51 kamila@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003"
* 18:51 kamila@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1262 - kamila@cumin1003"
* 18:46 kamila@cumin1003: START - Cookbook sre.dns.netbox
* 18:46 kamila@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1262
* 18:46 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 18:46 sukhe@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ncmonitor1001.eqiad.wmnet
* 18:46 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:45 kamila@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1262.eqiad.wmnet with OS trixie
* 18:45 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:45 kamila@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1262.eqiad.wmnet
* 18:44 kamila@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1262.eqiad.wmnet
* 18:44 kamila@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1262.eqiad.wmnet
* 18:42 sukhe@cumin1003: START - Cookbook sre.hosts.reboot-single for host ncmonitor1001.eqiad.wmnet
* 18:29 topranks: pull power on cr1-eqiad to install new switch-control boards [[phab:T426343|T426343]]
* 18:29 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[1018-1020].eqiad.wmnet with reason: line card install in cr1-eqiad
* 18:27 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 14 hosts with reason: linecard install in cr1-eqad
* 18:22 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_ulsfo
* 18:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4052.ulsfo.wmnet
* 18:19 cdobbins@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 18:15 cdobbins@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 18:14 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_drmrs
* 18:14 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6016.drmrs.wmnet
* 18:12 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_ulsfo
* 18:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4044.ulsfo.wmnet
* 18:10 sukhe@cumin1003: END (ERROR) - Cookbook sre.cdn.roll-reboot (exit_code=97) rolling reboot on A:cp-upload_drmrs
* 18:05 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2241: Security update
* 17:56 topranks: start draining traffic on cr1-eqiad ahead of line card installation [[phab:T426343|T426343]]
* 17:47 cdobbins@cumin2003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
* 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4051.ulsfo.wmnet
* 17:40 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:39 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 17:34 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6007.drmrs.wmnet
* 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6015.drmrs.wmnet
* 17:32 hnowlan@deploy1003: helmfile [eqiad] DONE helmfile.d/services/changeprop-jobqueue: apply
* 17:31 hnowlan@deploy1003: helmfile [eqiad] START helmfile.d/services/changeprop-jobqueue: apply
* 17:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4043.ulsfo.wmnet
* 17:27 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:25 lerickson@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 17:22 lerickson@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:wikikube-worker-codfw
* 17:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker-exp2001.codfw.wmnet
* 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 17:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker-exp2001.codfw.wmnet
* 17:22 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 17:19 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2241: Security update
* 17:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2241.codfw.wmnet
* 17:17 cwilliams@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2241.codfw.wmnet
* 17:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker-exp2001.codfw.wmnet
* 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker-exp2001.codfw.wmnet
* 17:15 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:15 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 17:10 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook-next: apply
* 17:08 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:06 sukhe: sre.dns.roll-reboot to resume later
* 17:06 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-reboot (exit_code=97) rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox)
* 17:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns5003.wikimedia.org
* 17:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2241: Security update
* 17:03 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2241: Security update
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-growthbook: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2366-2374].codfw.wmnet
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wmde: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-wikidata: apply
* 17:03 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2357-2365].codfw.wmnet
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 17:03 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2357-2365].codfw.wmnet
* 17:03 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-sre: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 16:57 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-search: apply
* 16:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2357-2365].codfw.wmnet
* 16:55 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 16:53 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6006.drmrs.wmnet
* 16:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 16:52 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6014.drmrs.wmnet
* 16:52 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 16:51 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2357-2365].codfw.wmnet
* 16:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4042.ulsfo.wmnet
* 16:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:50 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:49 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns5003.wikimedia.org
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-research: apply
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 16:44 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-platform-eng: apply
* 16:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4050.ulsfo.wmnet
* 16:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2347-2356].codfw.wmnet
* 16:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:dse-k8s-worker-codfw
* 16:35 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet
* 16:35 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet
* 16:34 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3004.wikimedia.org
* 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-ml: apply
* 16:33 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-main: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-fr-tech: apply
* 16:32 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 16:30 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-dev: apply
* 16:29 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet
* 16:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet
* 16:24 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2002.codfw.wmnet
* 16:24 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2002.codfw.wmnet
* 16:23 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3004.wikimedia.org
* 16:23 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2337-2346].codfw.wmnet
* 16:23 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:22 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:17 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2002.codfw.wmnet
* 16:13 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2002.codfw.wmnet
* 16:12 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2001.codfw.wmnet
* 16:12 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2001.codfw.wmnet
* 16:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1065.eqiad.wmnet with OS trixie
* 16:12 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6005.drmrs.wmnet
* 16:11 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6013.drmrs.wmnet
* 16:08 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4041.ulsfo.wmnet
* 16:08 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns3003.wikimedia.org
* 16:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2327-2336].codfw.wmnet
* 16:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2317-2326].codfw.wmnet
* 16:06 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2317-2326].codfw.wmnet
* 16:05 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2001.codfw.wmnet
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-test: apply
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 16:05 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-analytics-product: apply
* 16:03 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4049.ulsfo.wmnet
* 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2001.codfw.wmnet
* 16:00 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 16:00 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 16:00 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2063.codfw.wmnet with OS trixie
* 15:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2317-2326].codfw.wmnet
* 15:57 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns3003.wikimedia.org
* 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs-test2001.codfw.wmnet
* 15:54 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:54 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2317-2326].codfw.wmnet
* 15:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:49 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2004.codfw.wmnet
* 15:48 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:48 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage
* 15:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1065.eqiad.wmnet with reason: host reimage
* 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2003.codfw.wmnet
* 15:42 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:42 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:42 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2006.wikimedia.org
* 15:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage
* 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2307-2316].codfw.wmnet
* 15:37 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2002.codfw.wmnet
* 15:36 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:36 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2063.codfw.wmnet with reason: host reimage
* 15:31 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6004.drmrs.wmnet
* 15:31 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:31 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-wdqs2001.codfw.wmnet
* 15:31 btullis@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:dse-k8s-worker-codfw
* 15:30 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6012.drmrs.wmnet
* 15:28 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2006.wikimedia.org
* 15:27 btullis@cumin1003: END (PASS) - Cookbook sre.zookeeper.roll-restart-zookeeper (exit_code=0) for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons.
* 15:27 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:27 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4040.ulsfo.wmnet
* 15:24 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1065.eqiad.wmnet with OS trixie
* 15:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4048.ulsfo.wmnet
* 15:21 btullis@cumin1003: START - Cookbook sre.zookeeper.roll-restart-zookeeper for Zookeeper A:zookeeper-analytics cluster: Roll restart of jvm daemons.
* 15:21 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2297-2306].codfw.wmnet
* 15:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:20 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid public cluster: Reboot Druid nodes
* 15:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 15:17 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm
* 15:14 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2063.codfw.wmnet with OS trixie
* 15:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2005.wikimedia.org
* 15:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:11 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-eqiad
* 15:06 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be1064.eqiad.wmnet with OS trixie
* 15:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-be2062.codfw.wmnet with OS trixie
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2287-2296].codfw.wmnet
* 15:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2277-2286].codfw.wmnet
* 15:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2277-2286].codfw.wmnet
* 14:59 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2005.wikimedia.org
* 14:57 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-jumbo-eqiad
* 14:52 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-codfw
* 14:50 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6003.drmrs.wmnet
* 14:50 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:50 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host relforge1009.eqiad.wmnet
* 14:49 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2277-2286].codfw.wmnet
* 14:49 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6011.drmrs.wmnet
* 14:47 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 14:46 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 14:46 klausman@cumin1003: END (FAIL) - Cookbook sre.k8s.pool-depool-node (exit_code=99) pool for host ml-serve1001.eqiad.wmnet
* 14:46 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 14:45 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1001.eqiad.wmnet
* 14:45 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4039.ulsfo.wmnet
* 14:44 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1009.eqiad.wmnet
* 14:44 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-codfw
* 14:44 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns2004.wikimedia.org
* 14:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2277-2286].codfw.wmnet
* 14:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 14:41 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4047.ulsfo.wmnet
* 14:40 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1001.eqiad.wmnet
* 14:38 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage
* 14:36 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be2062.codfw.wmnet with reason: host reimage
* 14:36 topranks: disconnect power on cr2-eqiad to shut down device for switch fabric replacement [[phab:T426343|T426343]]
* 14:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 14:35 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1001.eqiad.wmnet
* 14:35 klausman@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>ml-serve1001.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 14:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 14:34 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-be1064.eqiad.wmnet with reason: host reimage
* 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:33 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:30 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns2004.wikimedia.org
* 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2267-2276].codfw.wmnet
* 14:29 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:29 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:24 btullis@cumin1003: END (PASS) - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas (exit_code=0) rolling reboot on A:schema-eqiad
* 14:22 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:20 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:20 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:19 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:18 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:17 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2257-2266].codfw.wmnet
* 14:16 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:16 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:16 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-be2062.codfw.wmnet with OS trixie
* 14:15 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-be1064.eqiad.wmnet with OS trixie
* 14:15 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1006.wikimedia.org
* 14:15 btullis@cumin1003: START - Cookbook sre.misc-clusters.roll-restart-reboot-eventschemas rolling reboot on A:schema-eqiad
* 14:14 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]]
* 14:11 jforrester@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:11 jforrester@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:10 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS bookworm
* 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6002.drmrs.wmnet
* 14:09 jforrester@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6010.drmrs.wmnet
* 14:06 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1006.wikimedia.org
* 14:06 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:05 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs2001.codfw.wmnet with OS bookworm
* 14:05 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4038.ulsfo.wmnet
* 14:02 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4046.ulsfo.wmnet
* 14:00 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2243,2248-2256].codfw.wmnet
* 14:00 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on cr1-eqiad with reason: switch upgrade and line card install
* 14:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:59 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:57 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:55 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-eqiad
* 13:55 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid public cluster: Reboot Druid nodes
* 13:53 topranks: switch routing-engine on cr2-eqiad resetting all interfaces [[phab:T417873|T417873]]
* 13:51 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1005.wikimedia.org
* 13:50 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:49 brouberol@cumin1003: END (PASS) - Cookbook sre.kafka.roll-restart-reboot-brokers (exit_code=0) rolling reboot on A:kafka-test-eqiad
* 13:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2207-2215,2242].codfw.wmnet
* 13:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1005.wikimedia.org
* 13:35 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:30 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2197-2206].codfw.wmnet
* 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6001.drmrs.wmnet
* 13:28 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp6009.drmrs.wmnet
* 13:28 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:28 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS bookworm
* 13:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4037.ulsfo.wmnet
* 13:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2001
* 13:22 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2001
* 13:22 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp4045.ulsfo.wmnet
* 13:21 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns1004.wikimedia.org
* 13:19 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[1018-1020].eqiad.wmnet with reason: switch upgrade and line card install
* 13:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2009.codfw.wmnet
* 13:18 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2001
* 13:18 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:17 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2001.codfw.wmnet 26.16.192.10.in-addr.arpa 6.2.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:17 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:17 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003"
* 13:17 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2001 - bking@cumin2003"
* 13:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2009.codfw.wmnet
* 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_drmrs
* 13:17 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:17 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_drmrs
* 13:17 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 15 hosts with reason: switch upgrade and line card install
* 13:17 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:15 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 13:13 bking@cumin2003: START - Cookbook sre.dns.netbox
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-jumbo-eqiad
* 13:13 brouberol@cumin1003: START - Cookbook sre.kafka.roll-restart-reboot-brokers rolling reboot on A:kafka-test-eqiad
* 13:13 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns1004.wikimedia.org
* 13:13 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and not (A:ulsfo or A:magru) and (A:dnsbox)
* 13:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_ulsfo
* 13:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:12 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_ulsfo
* 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2187-2196].codfw.wmnet
* 13:11 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 13:11 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 13:06 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 13:05 bking@cumin2003: END (FAIL) - Cookbook sre.wdqs.data-transfer (exit_code=99) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards
* 13:05 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2001
* 13:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2001.codfw.wmnet with OS bookworm
* 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:04 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2009.codfw.wmnet with OS trixie
* 13:03 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:03 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1001.eqiad.wmnet -> wcqs1002.eqiad.wmnet, repooling source-only afterwards
* 13:01 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 15s)
* 13:01 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 13:00 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 12:57 btullis@cumin1003: END (PASS) - Cookbook sre.druid.reboot-workers (exit_code=0) for Druid analytics cluster: Reboot Druid nodes
* 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2173-2179,2184-2186].codfw.wmnet
* 12:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:54 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:47 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:43 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage
* 12:41 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2163-2172].codfw.wmnet
* 12:41 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:40 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:40 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2009.codfw.wmnet with reason: host reimage
* 12:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:29 btullis@cumin1003: END (PASS) - Cookbook sre.ceph.roll-restart-reboot-server (exit_code=0) rolling reboot on A:cephosd-codfw
* 12:25 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2153-2162].codfw.wmnet
* 12:25 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:24 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2009
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2009
* 12:22 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2009
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:22 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2009.codfw.wmnet 139.0.192.10.in-addr.arpa 9.3.1.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:22 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003"
* 12:22 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2009 - mvernon@cumin2003"
* 12:16 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 12:15 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:15 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2009
* 12:15 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:15 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2009.codfw.wmnet with OS trixie
* 12:15 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 12:14 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 12:14 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:13 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 12:12 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2010.codfw.wmnet
* 12:11 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2010.codfw.wmnet
* 12:10 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2143-2152].codfw.wmnet
* 12:07 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2133-2142].codfw.wmnet
* 12:07 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2133-2142].codfw.wmnet
* 12:02 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 11:57 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2133-2142].codfw.wmnet
* 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2133-2142].codfw.wmnet
* 11:51 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:51 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:49 mvolz@deploy2003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:49 btullis@cumin1003: START - Cookbook sre.ceph.roll-restart-reboot-server rolling reboot on A:cephosd-codfw
* 11:48 mvolz@deploy2003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:47 btullis@cumin1003: START - Cookbook sre.druid.reboot-workers for Druid analytics cluster: Reboot Druid nodes
* 11:46 mvolz@deploy2003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:46 mvolz@deploy2003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:45 mvolz@deploy2003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:44 mvolz@deploy2003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:43 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:40 krinkle@deploy2003: Finished scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] (duration: 11m 38s)
* 11:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2010.codfw.wmnet with OS trixie
* 11:37 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2115,2124-2132].codfw.wmnet
* 11:36 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:36 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:36 krinkle@deploy2003: physikerwelt, krinkle: Continuing with deployment
* 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1018: Security updates
* 11:36 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:36 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:36 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1018: Security updates
* 11:31 krinkle@deploy2003: physikerwelt, krinkle: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:29 krinkle@deploy2003: Started scap sync-world: Backport for [[gerrit:1224074{{!}}Switch math rendering for group0 from native to mathjax (T413973)]]
* 11:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2105-2114].codfw.wmnet
* 11:20 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:20 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage
* 11:12 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2010.codfw.wmnet with reason: host reimage
* 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1018: Security updates
* 11:10 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 11:10 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 11:10 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1018: Security updates
* 11:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1009.eqiad.wmnet with OS trixie
* 11:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet
* 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2089-2095,2102-2104].codfw.wmnet
* 11:02 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 11:02 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=eqiad
* 11:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=eqiad
* 11:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet
* 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2010
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2010
* 10:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 10:54 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 10:54 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2010
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:54 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2010.codfw.wmnet 76.16.192.10.in-addr.arpa 6.7.0.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:54 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003"
* 10:54 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2010 - mvernon@cumin2003"
* 10:49 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2010
* 10:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage
* 10:49 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2010.codfw.wmnet with OS trixie
* 10:46 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2011.codfw.wmnet
* 10:45 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1011.eqiad.wmnet
* 10:44 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2011.codfw.wmnet
* 10:44 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2071-2078,2087-2088].codfw.wmnet
* 10:44 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 10:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1009.eqiad.wmnet with reason: host reimage
* 10:44 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:43 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow6001.drmrs.wmnet
* 10:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1017: Security updates
* 10:39 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:39 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:39 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1017: Security updates
* 10:39 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow6001.drmrs.wmnet
* 10:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1011.eqiad.wmnet
* 10:35 cgoubert@deploy2003: Finished deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]] (duration: 28m 34s)
* 10:35 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1012.eqiad.wmnet
* 10:35 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 10:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow5003.eqsin.wmnet
* 10:34 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2011.codfw.wmnet with OS trixie
* 10:33 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1009.eqiad.wmnet with OS trixie
* 10:28 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1012.eqiad.wmnet
* 10:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1013.eqiad.wmnet
* 10:27 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2059-2062,2064-2065,2067-2070].codfw.wmnet
* 10:27 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow5003.eqsin.wmnet
* 10:26 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:26 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow4003.ulsfo.wmnet
* 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:25 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow4003.ulsfo.wmnet
* 10:21 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1013.eqiad.wmnet
* 10:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps1014.eqiad.wmnet
* 10:16 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage
* 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1017: Security updates
* 10:14 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:14 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:14 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1017: Security updates
* 10:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow3004.esams.wmnet
* 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2011.codfw.wmnet with reason: host reimage
* 10:10 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2044,2046,2049-2051,2055-2058].codfw.wmnet
* 10:10 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps1014.eqiad.wmnet
* 10:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki2003.codfw.wmnet
* 10:09 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 10:09 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 10:09 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow3004.esams.wmnet
* 10:08 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2004.codfw.wmnet
* 10:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1010.eqiad.wmnet with OS trixie
* 10:07 cgoubert@deploy2003: Started deploy [restbase/deploy@06301bd]: Deploying {{Gerrit|1306088}} {{Gerrit|1308347}} - [[phab:T429944|T429944]] [[phab:T428279|T428279]]
* 10:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki2003.codfw.wmnet
* 10:04 topranks: push out config change to BGP_outfilter on core routers [[phab:T431849|T431849]]
* 10:02 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2004.codfw.wmnet
* 09:59 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 09:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow2003.codfw.wmnet
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2011
* 09:54 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2011
* 09:53 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2017-2018,2033-2039,2041].codfw.wmnet
* 09:52 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:52 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2011
* 09:52 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:52 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2011.codfw.wmnet 36.32.192.10.in-addr.arpa 6.3.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:51 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003"
* 09:51 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2011 - mvernon@cumin2003"
* 09:51 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow2003.codfw.wmnet
* 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1003.eqiad.wmnet
* 09:49 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 09:49 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 09:49 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage
* 09:47 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 09:47 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 09:47 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 09:47 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2011
* 09:46 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2011.codfw.wmnet with OS trixie
* 09:44 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1003.eqiad.wmnet
* 09:44 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2012.codfw.wmnet
* 09:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow1002.eqiad.wmnet
* 09:44 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1010.eqiad.wmnet with reason: host reimage
* 09:43 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2012.codfw.wmnet
* 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates
* 09:43 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:43 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:43 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates
* 09:42 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:40 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netflow1002.eqiad.wmnet
* 09:40 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 09:37 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host rpki1001.eqiad.wmnet
* 09:36 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 09:36 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 09:33 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host rpki1001.eqiad.wmnet
* 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2016].codfw.wmnet
* 09:32 cgoubert@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:wikikube-worker-codfw
* 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=eqiad
* 09:31 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=eqiad
* 09:31 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 09:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2012.codfw.wmnet with OS trixie
* 09:30 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1010.eqiad.wmnet with OS trixie
* 09:27 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1011.eqiad.wmnet with OS trixie
* 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates
* 09:21 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates
* 09:12 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage
* 09:08 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage
* 09:08 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2012.codfw.wmnet with reason: host reimage
* 09:05 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1011.eqiad.wmnet with reason: host reimage
* 08:55 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:54 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:52 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1011.eqiad.wmnet with OS trixie
* 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1022: Security updates
* 08:51 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2012
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2012
* 08:50 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:50 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1022: Security updates
* 08:50 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2012
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:50 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2012.codfw.wmnet 44.48.192.10.in-addr.arpa 4.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003"
* 08:50 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2012 - mvernon@cumin2003"
* 08:47 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1012.eqiad.wmnet with OS trixie
* 08:44 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 08:44 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2012
* 08:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2012.codfw.wmnet with OS trixie
* 08:42 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2013.codfw.wmnet
* 08:41 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2013.codfw.wmnet
* 08:35 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 08:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb1002.eqiad.wmnet
* 08:30 elukey@dns1004: END - running authdns-update
* 08:29 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage
* 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates
* 08:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:28 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:28 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates
* 08:27 elukey@dns1004: START - running authdns-update
* 08:26 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 08:26 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb1002.eqiad.wmnet
* 08:22 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1012.eqiad.wmnet with reason: host reimage
* 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host krb2002.codfw.wmnet
* 08:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast6003.wikimedia.org
* 08:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2013.codfw.wmnet with OS trixie
* 08:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast6003.wikimedia.org
* 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast3007.wikimedia.org
* 08:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host krb2002.codfw.wmnet
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1022: Security updates
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:09 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1022: Security updates
* 08:07 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1012.eqiad.wmnet with OS trixie
* 08:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast3007.wikimedia.org
* 08:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast5005.wikimedia.org
* 07:58 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast5005.wikimedia.org
* 07:58 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1013.eqiad.wmnet with OS trixie
* 07:53 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage
* 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates
* 07:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:53 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:53 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates
* 07:52 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast1004.wikimedia.org
* 07:47 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2013.codfw.wmnet with reason: host reimage
* 07:46 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host bast1004.wikimedia.org
* 07:40 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage
* 07:36 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1013.eqiad.wmnet with reason: host reimage
* 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2013
* 07:31 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2013
* 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates
* 07:30 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:30 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:30 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates
* 07:24 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2013
* 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:24 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2013.codfw.wmnet 87.0.192.10.in-addr.arpa 7.8.0.0.0.0.0.0.2.9.1.0.0.1.0.0.1.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 07:24 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003"
* 07:24 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2013 - mvernon@cumin2003"
* 07:20 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1013.eqiad.wmnet with OS trixie
* 07:19 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2013
* 07:19 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2013.codfw.wmnet with OS trixie
* 07:13 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] (duration: 07m 48s)
* 07:09 kharlan@deploy2003: kharlan: Continuing with deployment
* 07:08 kharlan@deploy2003: kharlan: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:06 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1310470{{!}}SiteStats: Add temporary account columns to site_stats (T339291)]]
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
* 01:15 lerickson@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 01:14 lerickson@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
== 2026-07-14 ==
* 22:51 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-upload_magru
* 22:51 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7016.magru.wmnet
* 22:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-reboot (exit_code=0) rolling reboot on A:cp-text_magru
* 22:46 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7008.magru.wmnet
* 22:09 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7015.magru.wmnet
* 22:04 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7007.magru.wmnet
* 21:29 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7014.magru.wmnet
* 21:23 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7006.magru.wmnet
* 21:13 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 0:15:00 on gerrit.wikimedia.org with reason: reboot
* 21:11 mutante: gerrit2003 (gerrit.wikimedia.org) - reboot for maintenance
* 21:11 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on gerrit2003.wikimedia.org with reason: reboot
* 20:56 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 20:56 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 20:56 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 20:55 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 20:48 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7013.magru.wmnet
* 20:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7005.magru.wmnet
* 20:41 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host phab1005.eqiad.wmnet with OS trixie
* 20:28 sbassett@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] (duration: 06m 47s)
* 20:24 sbassett@deploy2003: sbassett: Continuing with deployment
* 20:23 sbassett@deploy2003: sbassett: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:23 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage
* 20:21 sbassett@deploy2003: Started scap sync-world: Backport for [[gerrit:1310625{{!}}Add logging for various re-authentication methods (T432042)]]
* 20:20 aokoth@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on phab1005.eqiad.wmnet with reason: host reimage
* 20:12 jhuneidi@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] (duration: 07m 42s)
* 20:07 jhuneidi@deploy2003: jhuneidi, priyankar22: Continuing with deployment
* 20:06 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7012.magru.wmnet
* 20:06 jhuneidi@deploy2003: jhuneidi, priyankar22: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:04 jhuneidi@deploy2003: Started scap sync-world: Backport for [[gerrit:1309865{{!}}thwiki: Change to Wikipedia 25 logo (T431094)]]
* 20:02 aokoth@cumin1003: START - Cookbook sre.hosts.reimage for host phab1005.eqiad.wmnet with OS trixie
* 20:00 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7004.magru.wmnet
* 20:00 aokoth@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host vrts1004.eqiad.wmnet
* 19:57 aokoth@cumin1003: START - Cookbook sre.hosts.reboot-single for host vrts1004.eqiad.wmnet
* 19:24 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7011.magru.wmnet
* 19:19 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7003.magru.wmnet
* 19:11 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] (duration: 08m 33s)
* 19:07 jforrester@deploy2003: jforrester: Continuing with deployment
* 19:04 jforrester@deploy2003: jforrester: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 19:02 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310634{{!}}wikifunctions: Define the virtual-wikifunctions-usage table (T390557)]]
* 18:43 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7010.magru.wmnet
* 18:38 mutante: rotating phabricator-gerrit bot token (its-phabricator)
* 18:18 jhuneidi@deploy2003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 17:44 swfrench@deploy2003: Finished scap sync-world: Deployment to pick up new production image (duration: 31m 44s)
* 17:42 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7002.magru.wmnet
* 17:33 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7009.magru.wmnet
* 17:32 swfrench@deploy2003: swfrench: Continuing with deployment
* 17:29 swfrench@deploy2003: swfrench: Deployment to pick up new production image synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:17 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2035: repooling after rack b5 maintenance
* 17:16 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool es2035: repooling after rack b5 maintenance
* 17:16 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2188: repooling after rack b5 maintenance
* 17:12 swfrench@deploy2003: Started scap sync-world: Deployment to pick up new production image
* 17:01 sukhe@cumin1003: cookbooks.sre.cdn.roll-reboot finished rebooting cp7001.magru.wmnet
* 16:57 swfrench-wmf: reprepro include php8.3_8.3.32-1+wmf12u2 into component/php83 for bookworm-wikimedia
* 16:50 sukhe: pool cp2046
* 16:47 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=cp4039.ulsfo.wmnet
* 16:44 sukhe: sudo cumin -b31 "A:cp" "run-puppet-agent"
* 16:33 dzahn@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:15:00 on contint1003.wikimedia.org with reason: reboot
* 16:32 mutante: contint1003 - main CI server - rebooting
* 16:31 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2188: repooling after rack b5 maintenance
* 16:31 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2178: repooling after rack b5 maintenance
* 16:29 blake@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 16:28 blake@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 16:28 blake@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 16:28 blake@deploy2003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 16:18 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2014.codfw.wmnet
* 16:18 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 16:17 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2014.codfw.wmnet
* 16:10 mvernon@cumin1003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies (exit_code=0) rolling restart_daemons on A:thanos-fe
* 16:09 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 16:07 sukhe@puppetserver1001: conftool action : set/pooled=no; selector: name=cp4039.ulsfo.wmnet
* 16:07 mvernon@cumin1003: START - Cookbook sre.swift.roll-restart-reboot-swift-thanos-proxies rolling restart_daemons on A:thanos-fe
* 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2014.codfw.wmnet with OS trixie
* 15:56 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1014.eqiad.wmnet with OS trixie
* 15:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2014.codfw.wmnet with reason: host reimage
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ms-fe2014
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ms-fe2014
* 15:28 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host ms-fe2014
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:28 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache ms-fe2014.codfw.wmnet 194.16.192.10.in-addr.arpa 4.9.1.0.6.1.0.0.2.9.1.0.0.1.0.0.2.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:28 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003"
* 15:28 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ms-fe2014 - mvernon@cumin2003"
* 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003"
* 15:23 kamila@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003
* 15:22 kamila@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Apply title-related policies when selecting the name of the entity - kamila@cumin1003
* 15:22 kamila@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Apply title-related policies when selecting the name of the entity - kamila@cumin1003"
* 15:20 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host ms-fe2014
* 15:20 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2014.codfw.wmnet with OS trixie
* 15:19 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1014.eqiad.wmnet with OS trixie
* 15:01 dancy@deploy2003: Installation of scap version "4.274.1" completed for 3 hosts
* 15:00 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2177: repooling after rack b5 maintenance
* 15:00 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: repooling after rack b5 maintenance
* 14:59 dancy@deploy2003: Installing scap version "4.274.1" for 3 host(s)
* 14:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2015.codfw.wmnet with OS trixie
* 14:54 seanleong-wmde: Finished populateSitesTable for isvwiki ([[phab:T429939|T429939]])
* 14:53 javiermonton@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] (duration: 07m 35s)
* 14:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1015.eqiad.wmnet with OS trixie
* 14:49 javiermonton@deploy2003: javiermonton: Continuing with deployment
* 14:48 javiermonton@deploy2003: javiermonton: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:46 javiermonton@deploy2003: Started scap sync-world: Backport for [[gerrit:1310587{{!}}stream: pageview.v1 (T425624)]]
* 14:42 otto@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventstreams: apply
* 14:41 otto@deploy2003: helmfile [eqiad] START helmfile.d/services/eventstreams: apply
* 14:41 otto@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventstreams: apply
* 14:40 otto@deploy2003: helmfile [codfw] START helmfile.d/services/eventstreams: apply
* 14:40 otto@deploy2003: helmfile [staging] DONE helmfile.d/services/eventstreams: apply
* 14:39 otto@deploy2003: helmfile [staging] START helmfile.d/services/eventstreams: apply
* 14:38 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage
* 14:34 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage
* 14:33 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2015.codfw.wmnet with reason: host reimage
* 14:30 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1015.eqiad.wmnet with reason: host reimage
* 14:30 seanleong-wmde@deploy2003: mwscript-k8s job started: foreachwikiindblist wikidataclient extensions/Wikibase/lib/maintenance/populateSitesTable.php --force-protocol https # [[phab:T429939|T429939]]
* 14:24 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=0) rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum
* 14:16 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2015.codfw.wmnet with OS trixie
* 14:15 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2016.codfw.wmnet with OS trixie
* 14:14 cmooney@cumin1003: START - Cookbook sre.mysql.pool pool db2159: repooling after rack b5 maintenance
* 14:14 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1015.eqiad.wmnet with OS trixie
* 14:12 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1016.eqiad.wmnet with OS trixie
* 14:12 cmooney@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=pki,name=codfw
* 14:12 sbisson@deploy2003: helmfile [codfw] DONE helmfile.d/services/cxserver: sync
* 14:11 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve2002.codfw.wmnet
* 14:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve2002.codfw.wmnet
* 14:11 sbisson@deploy2003: helmfile [codfw] START helmfile.d/services/cxserver: sync
* 14:09 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1005.wikimedia.org
* 14:07 sbisson@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cxserver: sync
* 14:07 sbisson@deploy2003: helmfile [eqiad] START helmfile.d/services/cxserver: sync
* 14:05 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1005.wikimedia.org
* 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2005.wikimedia.org
* 14:02 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host dse-k8s-worker2003.codfw.wmnet
* 14:02 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host dse-k8s-worker2003.codfw.wmnet
* 14:01 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=tegola-vector-tiles,name=codfw
* 14:00 elukey@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=kartotherian,name=codfw
* 14:00 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2005.wikimedia.org
* 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage
* 13:57 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-ncredir (exit_code=0) rolling reboot on A:ncredir and A:ncredir
* 13:57 sbisson@deploy2003: helmfile [staging] DONE helmfile.d/services/cxserver: sync
* 13:56 sbisson@deploy2003: helmfile [staging] START helmfile.d/services/cxserver: sync
* 13:54 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage
* 13:52 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:52 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:51 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2016.codfw.wmnet with reason: host reimage
* 13:50 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1016.eqiad.wmnet with reason: host reimage
* 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy (exit_code=0) rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy
* 13:49 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling reboot on A:wikidough
* 13:46 sukhe@cumin1003: END (PASS) - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy (exit_code=0) rolling reboot on A:tcpproxy and A:tcpproxy
* 13:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and not (A:durum-eqiad or A:durum-codfw or A:durum-esams) and A:durum
* 13:42 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-durum (exit_code=97) rolling reboot on A:durum and A:durum
* 13:38 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2011.codfw.wmnet
* 13:36 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-reboot (exit_code=0) rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox)
* 13:36 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4004.wikimedia.org
* 13:34 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1016.eqiad.wmnet with OS trixie
* 13:34 topranks: reboot lsw1-b5-codfw to upgrade JunOS [[phab:T430918|T430918]]
* 13:34 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2016.codfw.wmnet with OS trixie
* 13:32 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve2002.codfw.wmnet
* 13:31 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2011.codfw.wmnet
* 13:31 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2012.codfw.wmnet
* 13:28 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2017.codfw.wmnet with OS trixie
* 13:25 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1017.eqiad.wmnet with OS trixie
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2012.codfw.wmnet
* 13:22 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve2002.codfw.wmnet
* 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:22 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 13:22 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4004.wikimedia.org
* 13:21 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2005.codfw.wmnet
* 13:19 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2013.codfw.wmnet
* 13:18 elukey@dns1004: END - running authdns-update
* 13:17 cmooney@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2005.codfw.wmnet
* 13:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2004.codfw.wmnet
* 13:16 elukey@dns1004: START - running authdns-update
* 13:16 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1046: es1046 after reimage
* 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s8
* 13:14 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1029.eqiad.wmnet,service=s5
* 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s5
* 13:13 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1029.eqiad.wmnet,service=s8
* 13:13 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2004.codfw.wmnet
* 13:12 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2013.codfw.wmnet
* 13:11 cmooney@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2095,2272-2278].codfw.wmnet
* 13:11 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps2014.codfw.wmnet
* 13:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aux-k8s-etcd2003.codfw.wmnet
* 13:09 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage
* 13:07 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.admin (exit_code=0) rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:07 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot finished rebooting dns4003.wikimedia.org
* 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host aux-k8s-etcd2003.codfw.wmnet
* 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1067.eqiad.wmnet
* 13:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1067.eqiad.wmnet
* 13:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1067.eqiad.wmnet
* 13:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idm-test1001.wikimedia.org
* 13:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2017.codfw.wmnet with reason: host reimage
* 13:05 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage
* 13:04 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps2014.codfw.wmnet
* 13:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.admin rebooting P<nowiki>{</nowiki>lvs7003*<nowiki>}</nowiki> and A:liberica
* 13:02 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2188: codfw rack B5 depool for maintenance
* 13:02 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-text_magru
* 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2188: codfw rack B5 depool for maintenance
* 13:01 sukhe@cumin1003: START - Cookbook sre.cdn.roll-reboot rolling reboot on A:cp-upload_magru
* 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2178: codfw rack B5 depool for maintenance
* 13:01 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idm-test1001.wikimedia.org
* 13:01 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2178: codfw rack B5 depool for maintenance
* 13:01 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2177: codfw rack B5 depool for maintenance
* 13:00 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2177: codfw rack B5 depool for maintenance
* 12:59 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1068.eqiad.wmnet
* 12:59 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1068.eqiad.wmnet
* 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola-vector-tiles,name=codfw
* 12:58 cmooney@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: codfw rack B5 depool for maintenance
* 12:58 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1017.eqiad.wmnet with reason: host reimage
* 12:58 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=tegola,name=codfw
* 12:57 elukey@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=kartotherian,name=codfw
* 12:57 cmooney@cumin1003: START - Cookbook sre.mysql.depool depool db2159: codfw rack B5 depool for maintenance
* 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 30 hosts with reason: lsw1-b5-codfw JunOS upgrade
* 12:55 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b5-codfw,lsw1-b5-codfw IPv6,lsw1-b5-codfw.mgmt,ssw1-a[1,8]-codfw.mgmt with reason: switch upgade lsw1-b5-codfw
* 12:54 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host maps-test2001.codfw.wmnet
* 12:49 cmooney@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=pki,name=codfw
* 12:49 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1067.eqiad.wmnet with OS trixie
* 12:48 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host maps-test2001.codfw.wmnet
* 12:48 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2017.codfw.wmnet with OS trixie
* 12:47 topranks: depool codfw pki in dns discovery ahead of lsw1-b5-codfw maintenance [[phab:T430918|T430918]]
* 12:47 sukhe@cumin1003: cookbooks.sre.dns.roll-reboot begin reboot of dns4003.wikimedia.org
* 12:47 sukhe@cumin1003: START - Cookbook sre.dns.roll-reboot rolling reboot on A:dnsbox and A:ulsfo and (A:dnsbox)
* 12:47 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=ms-fe2018.codfw.wmnet
* 12:45 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=ms-fe2018.codfw.wmnet
* 12:45 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-durum rolling reboot on A:durum and A:durum
* 12:45 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-tcp-proxy rolling reboot on A:tcpproxy and A:tcpproxy
* 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host pki-root1002.eqiad.wmnet
* 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 12:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1009.eqiad.wmnet
* 12:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1009.eqiad.wmnet
* 12:44 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-hcaptcha-proxy rolling reboot on A:hcaptcha-proxy and A:hcaptcha-proxy
* 12:43 sukhe@cumin1003: START - Cookbook sre.cdn.roll-restart-reboot-ncredir rolling reboot on A:ncredir and A:ncredir
* 12:43 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling reboot on A:wikidough
* 12:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2018.codfw.wmnet with OS trixie
* 12:42 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1017.eqiad.wmnet with OS trixie
* 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader2006.wikimedia.org
* 12:39 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1018.eqiad.wmnet with OS trixie
* 12:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1009.eqiad.wmnet
* 12:38 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host pki-root1002.eqiad.wmnet
* 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1009.eqiad.wmnet
* 12:38 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1008.eqiad.wmnet
* 12:38 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1008.eqiad.wmnet
* 12:35 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader2006.wikimedia.org
* 12:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host urldownloader1006.wikimedia.org
* 12:33 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1008.eqiad.wmnet
* 12:30 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1046: es1046 after reimage
* 12:29 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host es1046.eqiad.wmnet with OS trixie
* 12:29 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host urldownloader1006.wikimedia.org
* 12:28 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test2005.wikimedia.org
* 12:28 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1008.eqiad.wmnet
* 12:27 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1007.eqiad.wmnet
* 12:27 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1007.eqiad.wmnet
* 12:27 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage
* 12:25 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1068.eqiad.wmnet with reason: vacuum overlarge container dbs
* 12:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage
* 12:24 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test2005.wikimedia.org
* 12:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host idp-test1005.wikimedia.org
* 12:23 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model goodfaith --old ([[phab:T431159|T431159]])
* 12:22 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1007.eqiad.wmnet
* 12:22 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host idp-test1005.wikimedia.org
* 12:22 atsukoito: restarting pybal on lvs2013 `low-traffic` for https://gerrit.wikimedia.org/r/1310535
* 12:22 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1007.eqiad.wmnet
* 12:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1006.eqiad.wmnet
* 12:21 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1006.eqiad.wmnet
* 12:20 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage
* 12:19 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2018.codfw.wmnet with reason: host reimage
* 12:18 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1067.eqiad.wmnet with reason: host reimage
* 12:16 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1006.eqiad.wmnet
* 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1006.eqiad.wmnet
* 12:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1005.eqiad.wmnet
* 12:15 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1005.eqiad.wmnet
* 12:15 atsukoito: restarting pybal on lvs2014 for https://gerrit.wikimedia.org/r/1310535
* 12:12 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1018.eqiad.wmnet with reason: host reimage
* 12:11 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1005.eqiad.wmnet
* 12:11 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1005.eqiad.wmnet
* 12:10 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1004.eqiad.wmnet
* 12:10 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1004.eqiad.wmnet
* 12:09 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage
* 12:08 atsukoito: restarting pybal on lvs1019 `low-traffic` for https://gerrit.wikimedia.org/r/1310535
* 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1004.eqiad.wmnet
* 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1004.eqiad.wmnet
* 12:06 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1003.eqiad.wmnet
* 12:06 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1003.eqiad.wmnet
* 12:05 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on es1046.eqiad.wmnet with reason: host reimage
* 12:05 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003"
* 12:04 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "set ml-serve1001 back to active state - cmooney@cumin1003"
* 12:04 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1003.eqiad.wmnet
* 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1003.eqiad.wmnet
* 12:01 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker1002.eqiad.wmnet
* 12:01 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker1002.eqiad.wmnet
* 12:01 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 11:59 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2018.codfw.wmnet with OS trixie
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1067
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1067
* 11:59 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1067
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:59 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1067.eqiad.wmnet 17.48.64.10.in-addr.arpa 7.1.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:59 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003"
* 11:58 atsukoito: restarting pybal on lvs1018 `high-traffic2` for https://gerrit.wikimedia.org/r/1310535
* 11:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker1002.eqiad.wmnet
* 11:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2019.codfw.wmnet with OS trixie
* 11:56 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker1002.eqiad.wmnet
* 11:56 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 11:56 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1018.eqiad.wmnet with OS trixie
* 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 11:54 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:54 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1019.eqiad.wmnet with OS trixie
* 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1003.eqiad.wmnet
* 11:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:49 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:49 aikochou@deploy2003: helmfile [codfw] DONE helmfile.d/services/changeprop: sync
* 11:48 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host es1046.eqiad.wmnet with OS trixie
* 11:48 aikochou@deploy2003: helmfile [codfw] START helmfile.d/services/changeprop: sync
* 11:48 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310535
* 11:46 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es1046: Reimage to Trixie
* 11:44 marostegui@cumin1003: START - Cookbook sre.mysql.depool depool es1046: Reimage to Trixie
* 11:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5:00:00 on es1046.eqiad.wmnet with reason: Reimage to Trixie
* 11:42 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:42 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl1002.eqiad.wmnet
* 11:42 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl100*<nowiki>}</nowiki> and (A:aux-master-eqiad or A:aux-worker-eqiad)
* 11:42 aikochou@deploy2003: helmfile [eqiad] DONE helmfile.d/services/changeprop: sync
* 11:41 aikochou@deploy2003: helmfile [eqiad] START helmfile.d/services/changeprop: sync
* 11:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage
* 11:36 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] (duration: 09m 41s)
* 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:36 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:35 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:32 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage
* 11:32 jforrester@deploy2003: jforrester, gengh: Continuing with deployment
* 11:29 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2019.codfw.wmnet with reason: host reimage
* 11:28 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1019.eqiad.wmnet with reason: host reimage
* 11:28 jforrester@deploy2003: jforrester, gengh: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:26 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1310530{{!}}abstractwiki: Make cacheAbstractContentFragment throttling a global setting (T430898)]]
* 11:20 btullis@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host dse-k8s-worker2003.codfw.wmnet
* 11:20 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:19 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:15 btullis@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host dse-k8s-worker2003.codfw.wmnet
* 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 11:12 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1019.eqiad.wmnet with OS trixie
* 11:10 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2019.codfw.wmnet with OS trixie
* 11:10 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1020.eqiad.wmnet with OS trixie
* 11:10 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1067 - blake@cumin1003"
* 11:09 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] (duration: 12m 12s)
* 11:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2020.codfw.wmnet with OS trixie
* 11:03 kharlan@deploy2003: kharlan: Continuing with deployment
* 11:01 blake@cumin1003: START - Cookbook sre.dns.netbox
* 11:01 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:57 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307717{{!}}WikimediaAntiAbuse: Register wmgUse config and load extension (T431023)]]
* 10:55 kharlan@deploy2003: Finished scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] (duration: 31m 40s)
* 10:52 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage
* 10:52 marostegui@dns1004: START - running authdns-update
* 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage
* 10:49 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:48 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1020.eqiad.wmnet with reason: host reimage
* 10:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2020.codfw.wmnet with reason: host reimage
* 10:43 kharlan@deploy2003: kharlan: Continuing with deployment
* 10:42 kharlan@deploy2003: kharlan: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 10:32 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1020.eqiad.wmnet with OS trixie
* 10:29 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2159: Repooling after switchover
* 10:29 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1067
* 10:28 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1021.eqiad.wmnet with OS trixie
* 10:27 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1067.eqiad.wmnet with OS trixie
* 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:27 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1067.eqiad.wmnet
* 10:27 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1067.eqiad.wmnet
* 10:26 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1067.eqiad.wmnet
* 10:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2020.codfw.wmnet with OS trixie
* 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1055.eqiad.wmnet
* 10:26 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1055.eqiad.wmnet
* 10:26 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1055.eqiad.wmnet
* 10:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2021.codfw.wmnet with OS trixie
* 10:24 kharlan@deploy2003: Started scap sync-world: Backport for [[gerrit:1307716{{!}}extension-list: Add WikimediaAntiAbuse (T431023)]]
* 10:11 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1055.eqiad.wmnet with OS trixie
* 10:09 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage
* 10:05 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage
* 10:03 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129 revert
* 10:02 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1021.eqiad.wmnet with reason: host reimage
* 10:01 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2021.codfw.wmnet with reason: host reimage
* 09:58 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310129
* 09:50 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage
* 09:45 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1055.eqiad.wmnet with reason: host reimage
* 09:45 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1021.eqiad.wmnet with OS trixie
* 09:44 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1022.eqiad.wmnet with OS trixie
* 09:44 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2159: Repooling after switchover
* 09:44 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2021.codfw.wmnet with OS trixie
* 09:42 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2022.codfw.wmnet with OS trixie
* 09:31 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2159.codfw.wmnet
* 09:28 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1055
* 09:28 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1055
* 09:27 mvernon@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage
* 09:27 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1022.eqiad.wmnet with reason: host reimage
* 09:24 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage
* 09:21 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2022.codfw.wmnet with reason: host reimage
* 09:20 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2159: Rebooting db2159.codfw.wmnet
* 09:20 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2159: Rebooting db2159.codfw.wmnet
* 09:18 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 09:18 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 09:18 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 09:17 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2159.codfw.wmnet
* 09:13 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8] (duration: 02m 07s)
* 09:11 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (thin): Regular analytics weekly train THIN [analytics/refinery@ad6e05b8]
* 09:10 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1022.eqiad.wmnet with OS trixie
* 09:07 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1023.eqiad.wmnet with OS trixie
* 09:06 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8] (duration: 04m 49s)
* 09:04 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2022.codfw.wmnet with OS trixie
* 09:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2023.codfw.wmnet with OS trixie
* 09:01 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b]: Regular analytics weekly train [analytics/refinery@ad6e05b8]
* 09:01 a-pizzata@deploy2003: Finished deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8] (duration: 02m 01s)
* 09:00 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1055
* 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:00 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1055.eqiad.wmnet 50.32.64.10.in-addr.arpa 0.5.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:00 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:00 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003"
* 09:00 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1055 - blake@cumin1003"
* 08:59 a-pizzata@deploy2003: Started deploy [analytics/refinery@ad6e05b] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@ad6e05b8]
* 08:56 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2159 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94811 and previous config saved to /var/cache/conftool/dbconfig/20260714-085624-cwilliams.json
* 08:55 blake@cumin1003: START - Cookbook sre.dns.netbox
* 08:55 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1055
* 08:54 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1055.eqiad.wmnet with OS trixie
* 08:54 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1055.eqiad.wmnet
* 08:53 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1055.eqiad.wmnet
* 08:53 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1055.eqiad.wmnet
* 08:52 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2220 to s7 primary [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94810 and previous config saved to /var/cache/conftool/dbconfig/20260714-085239-cwilliams.json
* 08:51 cezmunsta: Starting s7 codfw failover from db2159 to db2220 - [[phab:T430920|T430920]]
* 08:48 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage
* 08:45 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2220 with weight 0 [[phab:T430920|T430920]]', diff saved to https://phabricator.wikimedia.org/P94809 and previous config saved to /var/cache/conftool/dbconfig/20260714-084553-cwilliams.json
* 08:45 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430920|T430920]]
* 08:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage
* 08:42 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1023.eqiad.wmnet with reason: host reimage
* 08:41 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2023.codfw.wmnet with reason: host reimage
* 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:34 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:29 marostegui@dns1004: END - running authdns-update
* 08:29 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netbox-dev2003.codfw.wmnet
* 08:27 marostegui@dns1004: START - running authdns-update
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 08:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2009.codfw.wmnet
* 08:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2009.codfw.wmnet
* 08:25 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1023.eqiad.wmnet with OS trixie
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.reboot-single for host netbox-dev2003.codfw.wmnet
* 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:24 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:24 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2023.codfw.wmnet with OS trixie
* 08:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1029.eqiad.wmnet with reason: reboot
* 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:21 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2009.codfw.wmnet
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:21 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2009.codfw.wmnet
* 08:20 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2008.codfw.wmnet
* 08:20 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2008.codfw.wmnet
* 08:15 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2008.codfw.wmnet
* 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2008.codfw.wmnet
* 08:14 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet
* 08:14 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet
* 08:13 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe1024.eqiad.wmnet with OS trixie
* 08:12 elukey@cumin1003: END (PASS) - Cookbook sre.pki.restart-reboot (exit_code=0) rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki)
* 08:12 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 08:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 08:10 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 08:09 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet
* 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet
* 08:08 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2006.codfw.wmnet
* 08:08 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2006.codfw.wmnet
* 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ms-fe2024.codfw.wmnet with OS trixie
* 08:03 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2006.codfw.wmnet
* 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2006.codfw.wmnet
* 08:02 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2005.codfw.wmnet
* 08:02 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2005.codfw.wmnet
* 07:58 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2005.codfw.wmnet
* 07:58 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2005.codfw.wmnet
* 07:57 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2004.codfw.wmnet
* 07:57 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2004.codfw.wmnet
* 07:54 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors
* 07:54 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors
* 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2004.codfw.wmnet
* 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2004.codfw.wmnet
* 07:53 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2003.codfw.wmnet
* 07:53 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2003.codfw.wmnet
* 07:53 mvernon@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage
* 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) pki.discovery.wmnet. on all recursors
* 07:49 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2003.codfw.wmnet
* 07:49 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache pki.discovery.wmnet. on all recursors
* 07:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage
* 07:48 elukey@cumin1003: START - Cookbook sre.pki.restart-reboot rolling reboot on P<nowiki>{</nowiki>pki*<nowiki>}</nowiki> and (A:pki)
* 07:46 mvernon@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe1024.eqiad.wmnet with reason: host reimage
* 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:45 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on ms-fe2024.codfw.wmnet with reason: host reimage
* 07:45 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative: apply
* 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2003.codfw.wmnet
* 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.misc-clusters.restart-reboot-config-master (exit_code=0) rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw)
* 07:44 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2002.codfw.wmnet
* 07:44 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2002.codfw.wmnet
* 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2002.codfw.wmnet
* 07:39 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors
* 07:39 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors
* 07:39 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2002.codfw.wmnet
* 07:39 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-worker2*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:36 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:35 elukey@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) config-master.discovery.wmnet. on all recursors
* 07:35 elukey@cumin1003: START - Cookbook sre.dns.wipe-cache config-master.discovery.wmnet. on all recursors
* 07:34 elukey@cumin1003: START - Cookbook sre.misc-clusters.restart-reboot-config-master rolling reboot on P<nowiki>{</nowiki>config-master*<nowiki>}</nowiki> and (A:config-master or A:config-master-eqiad or A:config-master-codfw)
* 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2003.codfw.wmnet
* 07:31 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:31 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:29 mvernon@cumin1003: START - Cookbook sre.hosts.reimage for host ms-fe1024.eqiad.wmnet with OS trixie
* 07:28 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host ms-fe2024.codfw.wmnet with OS trixie
* 07:26 elukey@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:26 elukey@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-ctrl2002.codfw.wmnet
* 07:26 elukey@cumin1003: START - Cookbook sre.k8s.reboot-nodes rolling reboot on P<nowiki>{</nowiki>aux-k8s-ctrl200*<nowiki>}</nowiki> and (A:aux-master-codfw or A:aux-worker-codfw)
* 07:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 07:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 06:50 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host lists1004.wikimedia.org
* 06:44 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host lists1004.wikimedia.org
* 06:25 marostegui@dns1004: END - running authdns-update
* 06:23 marostegui@dns1004: START - running authdns-update
* 06:22 marostegui@dns1004: END - running authdns-update
* 06:20 marostegui@dns1004: START - running authdns-update
* 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1026.eqiad.wmnet with reason: reboot
* 06:04 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync
* 06:04 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync
* 06:03 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 06:03 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 06:02 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
* 06:01 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
* 06:01 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
* 06:00 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
* 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:59 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:40 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:39 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 05:26 marostegui@dns1004: END - running authdns-update
* 05:24 marostegui@dns1004: START - running authdns-update
* 05:24 marostegui@dns1004: START - running authdns-update
* 05:13 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1004.wikimedia.org
* 05:07 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1004.wikimedia.org
* 04:01 mwpresync@deploy2003: Pruned MediaWiki: 1.47.0-wmf.8 (duration: 01m 07s)
* 03:39 mwpresync@deploy2003: Finished scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]] (duration: 36m 01s)
* 03:03 mwpresync@deploy2003: Started scap sync-world: testwikis to 1.47.0-wmf.11 refs [[phab:T430830|T430830]]
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 29s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-13 ==
* 23:33 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1064.eqiad.wmnet
* 23:33 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1064.eqiad.wmnet
* 23:08 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1064.eqiad.wmnet with reason: vacuum overlarge container dbs
* 23:06 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1069.eqiad.wmnet
* 23:06 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1069.eqiad.wmnet
* 22:34 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1069.eqiad.wmnet with reason: vacuum overlarge container dbs
* 21:18 maryum: Deployed security fix for [[phab:T321092|T321092]]
* 20:28 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb13u1 into main for trixie-wikimedia - [[phab:T424266|T424266]]
* 20:26 swfrench-wmf: reprepro include etcd-mirror_0.0.12-1+deb12u1 into main for bookworm-wikimedia - [[phab:T428495|T428495]]
* 20:23 dancy@deploy2003: Finished scap sync-world: Testing [[phab:T431635|T431635]] (duration: 03m 36s)
* 20:19 dancy@deploy2003: Started scap sync-world: Testing [[phab:T431635|T431635]]
* 20:18 dancy@deploy2003: Installation of scap version "4.274.0" completed for 3 hosts
* 20:16 dancy@deploy2003: Installing scap version "4.274.0" for 3 host(s)
* 20:12 kemayo@deploy2003: Finished scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] (duration: 08m 25s)
* 20:07 kemayo@deploy2003: soda, esanders, kemayo: Continuing with deployment
* 20:05 kemayo@deploy2003: soda, esanders, kemayo: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there
* 20:04 kemayo@deploy2003: Started scap sync-world: Backport for [[gerrit:1212157{{!}}Enable DiscussionTools visual enhancements on enwiki (T409297)]], [[gerrit:1310161{{!}}Disable mobile "exit the editor" survey phase 2 (T426135)]], [[gerrit:1304630{{!}}Add source tab to ukwikisource's "Архів" (Archive) namespace (T53980)]]
* 18:22 cwhite: lvextend vg0/srv +500g on centrallog hosts
* 18:19 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS trixie
* 17:46 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1071.eqiad.wmnet
* 17:46 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1071.eqiad.wmnet
* 17:13 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1071.eqiad.wmnet with reason: vacuum overlarge container dbs
* 17:07 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1065.eqiad.wmnet
* 17:07 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1065.eqiad.wmnet
* 17:06 dzahn@dns1006: END - running authdns-update
* 17:04 dzahn@dns1006: START - running authdns-update
* 17:01 dzahn@dns1006: END - running authdns-update
* 16:59 dzahn@dns1006: START - running authdns-update
* 16:51 dancy@deploy2003: Finished scap sync-world: testing [[phab:T428971|T428971]] (duration: 03m 37s)
* 16:47 dancy@deploy2003: Started scap sync-world: testing [[phab:T428971|T428971]]
* 16:45 atsukoito: restarting pybal on lvs1019 to flush IP address for `cirrussearch1122.eqiad.wmnet` after moving the vlan [[phab:T431311|T431311]]
* 16:42 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 16:42 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 16:42 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 16:42 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 16:42 Amir1: mwscript-k8s --follow --dblist=ores -- extensions/ORES/maintenance/PurgeScoreCache.php --model damaging --old ([[phab:T431159|T431159]])
* 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test
* 16:34 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 16:34 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test
* 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test
* 16:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 16:33 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test
* 16:31 dancy@deploy2003: Installation of scap version "4.273.0" completed for 159 hosts
* 16:29 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1065.eqiad.wmnet with reason: vacuum overlarge container dbs
* 16:27 dancy@deploy2003: Installing scap version "4.273.0" for 159 host(s)
* 16:27 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics-external: sync
* 16:27 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics-external: sync
* 16:26 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics-external: sync
* 16:26 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics-external: sync
* 16:22 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-analytics: sync
* 16:21 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-analytics: sync
* 16:21 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-analytics: sync
* 16:21 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-analytics: sync
* 16:19 javiermonton@deploy2003: helmfile [codfw] DONE helmfile.d/services/eventgate-main: sync
* 16:19 javiermonton@deploy2003: helmfile [codfw] START helmfile.d/services/eventgate-main: sync
* 16:18 javiermonton@deploy2003: helmfile [eqiad] DONE helmfile.d/services/eventgate-main: sync
* 16:17 javiermonton@deploy2003: helmfile [eqiad] START helmfile.d/services/eventgate-main: sync
* 15:59 atsukoito: restarting pybal on lvs1018 for https://gerrit.wikimedia.org/r/1310117
* 15:55 aikochou@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 15:50 atsukoito: restarting pybal on lvs1020 for https://gerrit.wikimedia.org/r/1310117
* 15:46 aikochou@deploy2003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revertrisk' for release 'main' .
* 15:43 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host kafka-logging1006.eqiad.wmnet
* 15:43 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host kafka-logging1006.eqiad.wmnet
* 15:41 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host ganeti-test[2001-2003].codfw.wmnet
* 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host ganeti-test[2001-2003].codfw.wmnet
* 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox1003.eqiad.wmnet
* 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox1003.eqiad.wmnet
* 15:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host netbox2003.codfw.wmnet
* 15:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host netbox2003.codfw.wmnet
* 15:36 sukhe: restart pybal on lvs1020
* 15:34 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet
* 15:34 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet
* 15:08 btullis@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 15:06 btullis@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:01 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:35 cdobbins@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 14:34 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 14:33 jgiannelos@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 14:32 jgiannelos@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 14:29 cdobbins@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 14:28 marostegui@dns1004: END - running authdns-update
* 14:28 btullis@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:27 marostegui@dns1004: START - running authdns-update
* 14:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1023.eqiad.wmnet with reason: reboot
* 14:18 btullis@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2009.codfw.wmnet with OS trixie
* 14:14 swfrench-wmf: start rolling run-puppet-agent on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]]
* 14:05 swfrench-wmf: disable-puppet on A:cp for ATS config change - [[phab:T428909|T428909]] [[phab:T431838|T431838]]
* 14:05 cdobbins@cumin2002: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
* 14:02 marostegui@dns1004: END - running authdns-update
* 14:00 marostegui@dns1004: START - running authdns-update
* 14:00 ladsgroup@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for ms-be1070.eqiad.wmnet
* 14:00 ladsgroup@cumin1003: START - Cookbook sre.hosts.remove-downtime for ms-be1070.eqiad.wmnet
* 13:58 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage
* 13:57 cdobbins@cumin2002: conftool action : set/pooled=no; selector: name=dns7002.*
* 13:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2009.codfw.wmnet with reason: host reimage
* 13:48 rscout@deploy2003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 13:48 rscout@deploy2003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 13:48 rscout@deploy2003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 13:47 rscout@deploy2003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 13:40 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 13:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2009.codfw.wmnet with OS trixie
* 13:30 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on ms-be1070.eqiad.wmnet with reason: vacuum overlarge container dbs
* 13:28 aude@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] (duration: 11m 12s)
* 13:23 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:22 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Continuing with deployment
* 13:22 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:19 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:19 aude@deploy2003: aikochou, javiermonton, aude, gkm563: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]] synced to the testservers
* 13:17 aude@deploy2003: Started scap sync-world: Backport for [[gerrit:1310059{{!}}Enable ChartWizard on the beta cluster (T431990)]], [[gerrit:1308656{{!}}streams: webrequest - pageview - trending (T430675)]], [[gerrit:1307438{{!}}EventStreamConfig: add page_revert_risk_wikidata_prediction_change.v1 (T420883)]], [[gerrit:1309894{{!}}Remove nonexistent autopatrolled group from Outreach Wiki (T431959)]]
* 13:01 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 13:01 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:00 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 12:59 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:52 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 12:51 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 12:48 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:48 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:47 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2008.codfw.wmnet with OS trixie
* 12:47 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:47 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:47 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:46 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 12:45 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:45 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 12:45 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 12:43 dreamyjazz@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] (duration: 07m 02s)
* 12:38 dreamyjazz@deploy2003: dreamyjazz: Continuing with deployment
* 12:37 dreamyjazz@deploy2003: dreamyjazz: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:36 dreamyjazz@deploy2003: Started scap sync-world: Backport for [[gerrit:1309684{{!}}ProductionServices: Drop NodeJS iPoid URL (T416623)]]
* 12:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage
* 12:23 Msz2001: Deployed changes to private code for Suggested Investigations
* 12:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2008.codfw.wmnet with reason: host reimage
* 12:20 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:19 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 12:17 atsuko@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 12:16 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:15 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] (duration: 07m 14s)
* 12:10 mszwarc@deploy2003: mszwarc: Continuing with deployment
* 12:09 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:07 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310058{{!}}Revert^2 "SI: Fix client side instrumentation" (T431977)]]
* 12:04 mszwarc@deploy2003: sync-world aborted: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]] (duration: 00m 29s)
* 12:03 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]]
* 12:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2008.codfw.wmnet with OS trixie
* 12:00 zabe@deploy2003: Finished scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] (duration: 07m 37s)
* 11:55 zabe@deploy2003: zabe: Continuing with deployment
* 11:54 zabe@deploy2003: zabe: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 11:52 zabe@deploy2003: Started scap sync-world: Backport for [[gerrit:1310064{{!}}Do not apply elwiki abusefilter settings to dewiki (T431934)]], [[gerrit:1310061{{!}}etcd: Add support for x4 (T431989)]]
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:43 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:35 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:34 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:33 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:30 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:28 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:27 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:17 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2007.codfw.wmnet with OS trixie
* 11:09 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1310027{{!}}SI: Fix client side instrumentation (T431977)]]
* 11:06 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s8
* 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=x3
* 11:00 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1016.eqiad.wmnet,service=s5
* 10:59 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage
* 10:52 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2007.codfw.wmnet with reason: host reimage
* 10:51 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host dse-k8s-worker1023
* 10:50 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host dse-k8s-worker1023
* 10:44 btullis@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host dse-k8s-worker1023
* 10:43 btullis@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host dse-k8s-worker1023
* 10:42 marostegui@cumin1003: dbctl commit (dc=all): 'Change x4 masters [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94804 and previous config saved to /var/cache/conftool/dbconfig/20260713-104248-marostegui.json
* 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:37 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 10:35 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 10:34 atsuko@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/opensearch-ttmserver: apply
* 10:33 marostegui@cumin1003: dbctl commit (dc=all): 'Push x4 initial dbctl config [[phab:T404715|T404715]]', diff saved to https://phabricator.wikimedia.org/P94803 and previous config saved to /var/cache/conftool/dbconfig/20260713-103259-marostegui.json
* 10:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2007.codfw.wmnet with OS trixie
* 09:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2006.codfw.wmnet with OS trixie
* 09:42 marostegui@dns1004: END - running authdns-update
* 09:40 marostegui@dns1004: START - running authdns-update
* 09:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage
* 09:22 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2006.codfw.wmnet with reason: host reimage
* 09:06 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1024.eqiad.wmnet with reason: reboot
* 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 09:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 09:01 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2006.codfw.wmnet with OS trixie
* 08:44 arthurtaylor@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:43 arthurtaylor@deploy2003: helmfile [codfw] START helmfile.d/services/wikidata-query-gui: apply
* 08:43 arthurtaylor@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:42 arthurtaylor@deploy2003: helmfile [eqiad] START helmfile.d/services/wikidata-query-gui: apply
* 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:42 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:41 arthurtaylor@deploy2003: helmfile [staging] DONE helmfile.d/services/wikidata-query-gui: apply
* 08:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2004.codfw.wmnet
* 08:38 arthurtaylor@deploy2003: helmfile [staging] START helmfile.d/services/wikidata-query-gui: apply
* 08:33 btullis@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host db1208.eqiad.wmnet
* 08:30 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=x3
* 08:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be2005.codfw.wmnet with OS trixie
* 08:28 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2004.codfw.wmnet
* 08:28 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup2003.codfw.wmnet
* 08:24 marostegui@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es1039: Repooling after testing
* 08:23 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on clouddb1016.eqiad.wmnet with reason: cloning
* 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s5
* 08:23 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1016.eqiad.wmnet,service=s8
* 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 08:21 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 08:17 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup2003.codfw.wmnet
* 08:17 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1004.eqiad.wmnet
* 08:14 btullis@cumin1003: START - Cookbook sre.hosts.reboot-single for host db1208.eqiad.wmnet
* 08:11 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host phab1005.eqiad.wmnet
* 08:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage
* 08:07 marostegui@dns1004: END - running authdns-update
* 08:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1004.eqiad.wmnet
* 08:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1003.eqiad.wmnet
* 08:07 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 08:06 trueg@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 08:05 marostegui@dns1004: START - running authdns-update
* 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be2005.codfw.wmnet with reason: host reimage
* 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 08:05 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host phab1005.eqiad.wmnet
* 08:05 marostegui@dns1004: START - running authdns-update
* 08:05 marostegui@dns1004: START - running authdns-update
* 08:05 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 08:04 marostegui@dns1004: START - running authdns-update
* 08:00 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit1003.wikimedia.org
* 07:58 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1003.eqiad.wmnet
* 07:58 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1002-dev.eqiad.wmnet
* 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 07:58 trueg@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 07:54 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1002-dev.eqiad.wmnet
* 07:54 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudbackup1001-dev.eqiad.wmnet
* 07:54 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit1003.wikimedia.org
* 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 07:53 marostegui@cumin1003: conftool action : set/weight=50; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:52 Msz2001: UTC morning backport+config window done
* 07:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be2005.codfw.wmnet with OS trixie
* {{safesubst:SAL entry|1=07:50 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943}}
* 07:49 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudbackup1001-dev.eqiad.wmnet
* 07:46 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:46 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 07:45 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s6
* 07:44 marostegui@cumin1003: conftool action : set/weight=10; selector: name=clouddb1028.eqiad.wmnet,service=s4
* 07:43 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Continuing with deployment
* {{safesubst:SAL entry|1=07:39 mszwarc@deploy2003: mszwarc, danielyepezgarces, anzx: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark}}
* 07:39 marostegui@cumin1003: START - Cookbook sre.mysql.pool pool es1039: Repooling after testing
* {{safesubst:SAL entry|1=07:36 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309302{{!}}Enable campaignEvents on cowikimedia (T431765)]], [[gerrit:1307836{{!}}extwiki: Rename wgSitename to Güiquipedia and remove obsolete namespace alias (T431334)]], [[gerrit:1309901{{!}}cowikimedia: Update localtimezone to America/Bogota (T431965)]], [[gerrit:1308652{{!}}minwikiquote: set sitename, timezone and projectnamespace & add logo, wordmark (T429943)}}
* 07:35 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] (duration: 30m 03s)
* 07:25 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gerrit2002.wikimedia.org
* 07:22 mszwarc@deploy2003: mszwarc: Continuing with deployment
* 07:21 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:19 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gerrit2002.wikimedia.org
* 07:15 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host aphlict1002.eqiad.wmnet
* 07:11 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host aphlict1002.eqiad.wmnet
* 07:08 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2003.wikimedia.org
* 07:05 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309559{{!}}plwiki: Switch back to normal tagline (T430512)]]
* 07:02 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2003.wikimedia.org
* 07:02 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab2002.wikimedia.org
* 06:55 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab2002.wikimedia.org
* 06:55 jelto@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab1003.wikimedia.org
* 06:49 jelto@cumin1003: START - Cookbook sre.hosts.reboot-single for host gitlab1003.wikimedia.org
* 06:34 marostegui: Drop m5 ipoid database [[phab:T431007|T431007]]
* 06:29 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1027.eqiad.wmnet with reason: reboot
* 06:24 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1028.eqiad.wmnet with reason: reboot
* 06:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1025.eqiad.wmnet with reason: reboot
* 06:17 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dbproxy1022.eqiad.wmnet with reason: reboot
* 06:03 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on dbproxy[2005-2008].codfw.wmnet with reason: reboot
* 05:37 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db[1217,1228].eqiad.wmnet with reason: cloning
* 05:11 marostegui: Drop users_to_rename table [[phab:T431842|T431842]]
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 37s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-12 ==
* 16:01 marostegui@cumin1003: dbctl commit (dc=all): 'Depool db2209 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94792 and previous config saved to /var/cache/conftool/dbconfig/20260712-160124-marostegui.json
* 15:58 marostegui@cumin1003: dbctl commit (dc=all): 'Promote db2205 to s3 primary [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94791 and previous config saved to /var/cache/conftool/dbconfig/20260712-155853-marostegui.json
* 15:58 marostegui: Starting s3 codfw emergency failover from db2209 to db2205 - [[phab:T431950|T431950]]
* 15:51 marostegui@cumin1003: dbctl commit (dc=all): 'Set db2205 with weight 0 [[phab:T431950|T431950]]', diff saved to https://phabricator.wikimedia.org/P94790 and previous config saved to /var/cache/conftool/dbconfig/20260712-155135-marostegui.json
* 15:51 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T431950|T431950]]
* 02:01 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 01m 17s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-11 ==
* 02:06 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 26s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-10 ==
* 19:12 jhathaway@dns1004: END - running authdns-update
* 19:10 jhathaway@dns1004: START - running authdns-update
* 18:23 mutante: vrts2002 rebooting (not the active host)
* 18:21 mutante: lists2001, phab2003 - rebooting (not the active hosts)
* 18:16 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on A:lvs-high-traffic2-codfw
* 18:15 swfrench@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on A:lvs-high-traffic2-codfw
* 17:15 mutante: [doc1004:~] $ sudo systemctl start rsync-doc-host-data-sync ([[phab:T431856|T431856]])
* 17:09 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1004.eqiad.wmnet
* 17:08 jhathaway@dns1004: END - running authdns-update
* 17:07 jhathaway@dns1004: START - running authdns-update
* 17:06 jhathaway: depooling puppetserver1002, cause of errors is still unknown
* 17:03 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1004.eqiad.wmnet
* 16:57 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner1003.eqiad.wmnet
* 16:51 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner1003.eqiad.wmnet
* 16:48 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 16:48 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2004.codfw.wmnet
* 16:42 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2004.codfw.wmnet
* 16:41 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2003.codfw.wmnet
* 16:35 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2003.codfw.wmnet
* 16:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host gitlab-runner2002.codfw.wmnet
* 16:27 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-single for host gitlab-runner2002.codfw.wmnet
* 16:25 mutante: gitlab-runners (production) rebooting cluster one by one
* 16:17 mutante: etherpad1004/etherpad2002 - (etherpad.wikimedia.org) - rebooting
* 16:13 mutante: doc1004/doc2003 (doc.wikimedia.org backends) - rebooting
* 16:02 mutante: releases1003/releases2003 (releases.wikimedia.org backends) - rebooting for maintenance
* 15:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2007-dev.codfw.wmnet
* 15:19 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet
* 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host cloudcephosd2007-dev.codfw.wmnet
* 15:14 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2007-dev.codfw.wmnet
* 15:14 filippo@cumin1003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host cloudcephosd2006-dev.codfw.wmnet
* 15:07 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2006-dev.codfw.wmnet
* 15:07 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2005-dev.codfw.wmnet
* 14:59 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2005-dev.codfw.wmnet
* 14:59 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephosd2004-dev.codfw.wmnet
* 14:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephosd2004-dev.codfw.wmnet
* 14:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2007-dev.codfw.wmnet
* 14:51 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1054.eqiad.wmnet
* 14:51 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1054.eqiad.wmnet
* 14:51 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1054.eqiad.wmnet
* 14:47 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2007-dev.codfw.wmnet
* 14:47 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2006-dev.codfw.wmnet
* 14:41 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2006-dev.codfw.wmnet
* 14:40 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcephmon2005-dev.codfw.wmnet
* 14:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcephmon2005-dev.codfw.wmnet
* 14:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2005-dev.codfw.wmnet
* 14:29 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2005-dev.codfw.wmnet
* 14:29 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2006-dev.codfw.wmnet
* 14:21 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2006-dev.codfw.wmnet
* 14:21 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudcontrol2010-dev.codfw.wmnet
* 14:15 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudcontrol2010-dev.codfw.wmnet
* 14:15 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2004-dev.codfw.wmnet
* 14:10 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1054.eqiad.wmnet with OS trixie
* 14:09 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2004-dev.codfw.wmnet
* 14:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudgw2003-dev.codfw.wmnet
* 14:02 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudgw2003-dev.codfw.wmnet
* 14:02 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2004-dev.codfw.wmnet
* 13:53 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2004-dev.codfw.wmnet
* 13:53 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2003-dev.codfw.wmnet
* 13:48 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:44 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2003-dev.codfw.wmnet
* 13:44 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudlb2002-dev.codfw.wmnet
* 13:42 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 13:41 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 13:37 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudlb2002-dev.codfw.wmnet
* 13:37 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudidp2001-dev.codfw.wmnet
* 13:33 blake@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage
* 13:33 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudidp2001-dev.codfw.wmnet
* 13:33 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2006-dev.codfw.wmnet
* 13:26 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2006-dev.codfw.wmnet
* 13:26 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudnet2005-dev.codfw.wmnet
* 13:23 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1054.eqiad.wmnet with reason: host reimage
* 13:18 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudnet2005-dev.codfw.wmnet
* 13:18 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2005-dev.codfw.wmnet
* 13:12 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2005-dev.codfw.wmnet
* 13:11 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudservices2004-dev.codfw.wmnet
* 13:08 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudservices2004-dev.codfw.wmnet
* 13:08 filippo@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudweb2002-dev.wikimedia.org
* 13:05 jgiannelos@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 13:05 jgiannelos@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1054
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1054
* 13:04 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1054
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:04 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1054.eqiad.wmnet 49.32.64.10.in-addr.arpa 9.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:04 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003"
* 13:04 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1054 - blake@cumin1003"
* 13:01 filippo@cumin1003: START - Cookbook sre.hosts.reboot-single for host cloudweb2002-dev.wikimedia.org
* 13:00 blake@cumin1003: START - Cookbook sre.dns.netbox
* 12:59 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1054
* 12:57 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1054.eqiad.wmnet with OS trixie
* 12:57 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1054.eqiad.wmnet
* 12:56 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1054.eqiad.wmnet
* 12:56 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1054.eqiad.wmnet
* 12:47 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wcqs1002.eqiad.wmnet with OS trixie
* 12:44 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 12:39 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 12:38 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:37 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:14 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:10 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:09 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:08 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 12:07 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 12:00 gmodena@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/services/wdqs: apply
* 11:51 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:49 gmodena@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/services/wdqs: apply
* 11:48 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:47 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:44 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:32 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2001.codfw.wmnet
* 11:32 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1053.eqiad.wmnet
* 11:32 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2001.codfw.wmnet
* 11:32 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1053.eqiad.wmnet
* 11:32 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1053.eqiad.wmnet
* 11:31 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker2001.codfw.wmnet
* 11:31 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker2001.codfw.wmnet
* 11:31 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw)
* 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:30 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:21 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 18 hosts with reason: reboot & upgrade
* 11:20 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS trixie
* 11:16 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:15 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:14 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 11:14 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 11 hosts
* 11:14 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 11 hosts
* 11:13 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 11:08 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 11:02 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1053.eqiad.wmnet with OS trixie
* 11:01 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:58 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 10:57 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:38 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS trixie
* 10:38 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker2001.codfw.wmnet
* 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker2001.codfw.wmnet
* 10:38 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>wikikube-worker2001*<nowiki>}</nowiki> and (A:wikikube-master-codfw or A:wikikube-worker-codfw)
* 10:35 topranks: adjust IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]] towards ssw1-e1-codfw
* 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 cgoubert@cumin1003: START - Cookbook sre.hosts.remove-downtime for wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2002,2005-2006,2011-2012].codfw.wmnet
* 10:30 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1004.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=1) rolling reimage on A:wikikube-worker-codfw
* 10:27 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host wikikube-worker2001.codfw.wmnet with OS bookworm
* 10:25 cgoubert@cumin1003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:24 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:15 cgoubert@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker2001.codfw.wmnet with reason: host reimage
* 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 10:11 brouberol@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 10:08 gmodena@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/wdqs: apply
* 10:07 brouberol@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 10:06 brouberol@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 10:00 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:55 cgoubert@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker2001.codfw.wmnet with OS bookworm
* 09:55 cgoubert@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet
* 09:55 gmodena@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/services/wdqs: apply
* 09:52 cgoubert@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2001-2002,2005-2006,2011-2012].codfw.wmnet
* 09:51 cgoubert@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on A:wikikube-worker-codfw
* 09:41 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage
* 09:37 topranks: apply new IBGP outbound policy on lsw1-e2-codfw [[phab:T423430|T423430]]
* 09:36 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:36 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1053.eqiad.wmnet with reason: host reimage
* 09:16 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1053
* 09:16 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1053
* 09:15 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1053
* 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:15 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1053.eqiad.wmnet 48.32.64.10.in-addr.arpa 8.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:15 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:15 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003"
* 09:15 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1053 - blake@cumin1003"
* 09:11 blake@cumin1003: START - Cookbook sre.dns.netbox
* 09:11 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1053
* 09:08 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1053.eqiad.wmnet with OS trixie
* 09:08 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet
* 09:08 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet
* 09:08 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1053.eqiad.wmnet
* 09:06 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1053.eqiad.wmnet
* 09:04 brouberol@dns1004: END - running authdns-update
* 09:03 brouberol@dns1004: START - running authdns-update
* 08:41 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea] (duration: 02m 11s)
* 08:38 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (thin): Regular analytics weekly train THIN [analytics/refinery@1abf22ea]
* 08:38 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea] (duration: 05m 17s)
* 08:38 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:34 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:33 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e]: Regular analytics weekly train [analytics/refinery@1abf22ea]
* 08:32 javiermonton@deploy2003: Finished deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea] (duration: 02m 03s)
* 08:30 javiermonton@deploy2003: Started deploy [analytics/refinery@1abf22e] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@1abf22ea]
* 08:30 JavierMonton: Deploying Refinery at {{Gerrit|1abf22ea}} for changes 1308121/T427068 1306491/T430020 and {{Gerrit|1308190}}
* 08:29 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:29 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:24 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:24 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:18 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 08:18 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 08:00 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db[2183-2184].codfw.wmnet
* 08:00 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for db[2183-2184].codfw.wmnet
* 07:52 dcausse@deploy2003: helmfile [eqiad] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:52 dcausse@deploy2003: helmfile [eqiad] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 11 hosts with reason: reboot & upgrade
* 07:47 dcausse@deploy2003: helmfile [codfw] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:47 dcausse@deploy2003: helmfile [codfw] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:44 dcausse@deploy2003: helmfile [staging] DONE helmfile.d/services/cirrus-streaming-updater: apply
* 07:44 dcausse@deploy2003: helmfile [staging] START helmfile.d/services/cirrus-streaming-updater: apply
* 07:23 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 10 hosts
* 07:23 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 10 hosts
* 06:45 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 10 hosts with reason: reboot & upgrade
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 41s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-09 ==
* 23:33 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] (duration: 13m 26s)
* 23:29 ladsgroup@deploy2003: ladsgroup, jdlrobson: Continuing with deployment
* 23:22 ladsgroup@deploy2003: ladsgroup, jdlrobson: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 23:20 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1308671{{!}}Enable section share on Persian Wikipedia (T431514)]]
* 22:57 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.renumber-node (exit_code=0) Renumbering for host wikikube-worker1165.eqiad.wmnet
* 22:56 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1165.eqiad.wmnet
* 22:56 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1165.eqiad.wmnet
* 22:45 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1165.eqiad.wmnet with OS trixie
* 22:38 arlolra@deploy2003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 22:37 arlolra@deploy2003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 22:37 arlolra@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 22:37 arlolra@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:37 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 22:25 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage
* 22:17 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1165.eqiad.wmnet with reason: host reimage
* 22:13 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 22:12 rzl@deploy2003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 22:04 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 22:04 rzl@deploy2003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1165
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1165
* 22:02 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1165
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:02 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1165.eqiad.wmnet 115.48.64.10.in-addr.arpa 5.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 22:02 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002"
* 22:02 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1165 - jasmine@cumin2002"
* 22:02 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 21:57 jasmine@cumin2002: START - Cookbook sre.dns.netbox
* 21:55 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1165
* 21:54 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1165.eqiad.wmnet with OS trixie
* 21:54 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 21:54 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1165.eqiad.wmnet
* 21:53 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1165.eqiad.wmnet
* 21:53 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1165.eqiad.wmnet
* 21:53 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 21:47 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 21:45 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 21:43 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 21:43 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:42 maryum: Deploy fix for [[phab:T431684|T431684]]
* 21:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1002.eqiad.wmnet with reason: host reimage
* 21:27 ladsgroup@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] (duration: 34m 14s)
* 21:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1002
* 21:23 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1002
* 21:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1002.eqiad.wmnet with OS trixie
* 21:22 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 21:20 bking@deploy2003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 22s)
* 21:20 bking@deploy2003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 21:16 bking@cumin2003: END (ERROR) - Cookbook sre.wdqs.data-transfer (exit_code=97) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 21:16 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2001.codfw.wmnet -> wcqs2002.codfw.wmnet, repooling source-only afterwards
* 21:15 ladsgroup@deploy2003: ladsgroup: Continuing with deployment
* 21:13 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2002.codfw.wmnet with OS bookworm
* 21:11 ladsgroup@deploy2003: ladsgroup: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 21:08 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0)
* 21:07 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 6 hosts with reason: reboots
* 20:54 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003
* 20:53 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:53 ladsgroup@deploy2003: Started scap sync-world: Backport for [[gerrit:1309290{{!}}Introduce sharing of section function (T18691)]], [[gerrit:1309291{{!}}Drop the share icon from the mobile site (T18691)]]
* 20:51 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99)
* 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 20:47 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecycle work - bking@cumin2003
* 20:42 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 20:41 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:41 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99)
* 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=1) for host relforge1008.eqiad.wmnet
* 20:40 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1009.eqiad.wmnet with OS trixie
* 20:33 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:32 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:32 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster cloudelastic: lifecyle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:31 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:31 ladsgroup@cumin1003: END (PASS) - Cookbook sre.wikireplicas.update-views (exit_code=0)
* 20:29 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet
* 20:24 rzl@deploy2003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 20:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS bookworm
* 20:23 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host relforge1008.eqiad.wmnet
* 20:23 rzl@deploy2003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 20:23 bking@cumin2003: START - Cookbook sre.hosts.reboot-single for host relforge1008.eqiad.wmnet
* 20:22 bking@cumin2003: END (FAIL) - Cookbook sre.elasticsearch.rolling-operation (exit_code=99) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:22 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:21 bking@cumin2003: END (ERROR) - Cookbook sre.elasticsearch.rolling-operation (exit_code=97) Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:21 bking@cumin2003: START - Cookbook sre.elasticsearch.rolling-operation Operation.REBOOT (1 nodes at a time) for ElasticSearch cluster relforge: Lifecycle work - bking@cumin2003 - [[phab:T431658|T431658]]
* 20:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage
* 20:16 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 20:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1009.eqiad.wmnet with reason: host reimage
* 20:12 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.wikireplicas.update-views (exit_code=99)
* 20:02 ladsgroup@cumin1003: START - Cookbook sre.wikireplicas.update-views
* 19:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1009.eqiad.wmnet with OS trixie
* 19:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:43 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1009.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 19:30 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 19:28 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 19:27 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 19:25 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:41 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 17:45 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for doh5004.wikimedia.org
* 17:45 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for doh5004.wikimedia.org
* 17:38 ladsgroup@deploy2003: helmfile [eqiad] DONE helmfile.d/services/thumbor: apply
* 17:35 ladsgroup@deploy2003: helmfile [eqiad] START helmfile.d/services/thumbor: apply
* 17:29 ladsgroup@deploy2003: helmfile [codfw] DONE helmfile.d/services/thumbor: apply
* 17:26 ladsgroup@deploy2003: helmfile [codfw] START helmfile.d/services/thumbor: apply
* 17:09 mutante: zuul[12]00[123] - rebooting for maintenance
* 17:09 ebernhardson: start full in-place reindex of eqiad cirrussearch cluster
* 17:08 dzahn@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-cluster (exit_code=99)
* 17:08 dzahn@cumin2002: START - Cookbook sre.hosts.reboot-cluster
* 17:03 ebernhardson: start full in-place reindex of codfw cirrussearch cluster
* 16:59 mutante: stewards1001/stewards2001 - reboot for maintenance
* 16:54 ebernhardson: start full in-place reindex of cloudelastic cluster
* 16:53 ladsgroup@deploy2003: helmfile [staging] DONE helmfile.d/services/thumbor: apply
* 16:52 ladsgroup@deploy2003: helmfile [staging] START helmfile.d/services/thumbor: apply
* 16:49 mutante: planet1003/planet2003 - rebooting
* 16:47 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on doh5004.wikimedia.org with reason: random high load, investigating
* 15:55 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 15:54 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 15:51 jynus: restarting backupmon1001
* 15:49 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts
* 15:49 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts
* 15:47 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backupmon1001.eqiad.wmnet with reason: restart
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:13 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 15:06 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen-next: apply
* 14:59 cjming@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:58 cjming@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/test-kitchen: apply
* 14:51 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 14 hosts
* 14:51 jynus@cumin1003: START - Cookbook sre.hosts.remove-downtime for 14 hosts
* 14:49 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade
* 14:48 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1002.eqiad.wmnet
* 14:48 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1002.eqiad.wmnet
* 14:47 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:45 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1052.eqiad.wmnet
* 14:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1052.eqiad.wmnet
* 14:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1052.eqiad.wmnet
* 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:32 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:31 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:31 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:30 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:30 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/turnilo: sync
* 14:28 elukey@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/turnilo: sync
* 14:26 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:23 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:20 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1052.eqiad.wmnet with OS trixie
* 14:19 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on 6 hosts with reason: reboot & upgrade
* 14:18 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:17 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:15 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:14 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:13 elukey: update druid indexation job for webrequest_sampled_live - [[phab:T427068|T427068]]
* 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:11 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:09 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002"
* 14:09 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: updating for papaul - jhancock@cumin2002"
* 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:07 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/webrequest-page-view-next: apply
* 14:04 jhancock@cumin2002: START - Cookbook sre.dns.netbox
* 14:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cuminunpriv1001.eqiad.wmnet
* 13:59 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb1003.eqiad.wmnet
* 13:59 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage
* 13:57 moritzm: installing requests security updates
* 13:56 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cuminunpriv1001.eqiad.wmnet
* 13:55 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb1003.eqiad.wmnet
* 13:53 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1052.eqiad.wmnet with reason: host reimage
* 13:50 moritzm: installing python-cryptography security updates
* 13:47 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netboxdb2003.codfw.wmnet
* 13:44 Msz2001: UTC afternoon config+backport window is done
* 13:44 Msz2001: Updated `logging` on `metawiki` to fix log performers, [[phab:T431176|T431176]]#12105297
* 13:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netboxdb2003.codfw.wmnet
* 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:43 javiermonton@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt1002.wikimedia.org
* 13:41 mszwarc@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] (duration: 07m 30s)
* 13:41 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt1002.wikimedia.org
* 13:37 mszwarc@deploy2003: mszwarc: Continuing with deployment
* 13:36 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1052
* 13:36 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1052
* 13:35 mszwarc@deploy2003: mszwarc: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:35 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1052
* 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:35 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1052.eqiad.wmnet 47.32.64.10.in-addr.arpa 7.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:35 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:35 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003"
* 13:35 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1052 - blake@cumin1003"
* 13:34 mszwarc@deploy2003: Started scap sync-world: Backport for [[gerrit:1309101{{!}}Fix UserEditTracker spoiling ActorStore cache for cross-wiki lookups (T431176)]], [[gerrit:1309122{{!}}SuggestedInvestigations: Instrument link clicks in the cases table (T429320)]]
* 13:31 blake@cumin1003: START - Cookbook sre.dns.netbox
* 13:31 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1052
* 13:30 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1052.eqiad.wmnet with OS trixie
* 13:30 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1052.eqiad.wmnet
* 13:29 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1052.eqiad.wmnet
* 13:29 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1052.eqiad.wmnet
* 13:17 jforrester@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] (duration: 11m 26s)
* 13:13 jforrester@deploy2003: jforrester: Continuing with deployment
* 13:08 jforrester@deploy2003: jforrester: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:06 jforrester@deploy2003: Started scap sync-world: Backport for [[gerrit:1309112{{!}}abstractwiki: Don't run our jobs too quickly, they'll collide and error (T430898)]]
* 12:54 cgoubert@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mobileapps: apply
* 12:54 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:52 cgoubert@deploy2003: helmfile [eqiad] START helmfile.d/services/mobileapps: apply
* 12:45 cgoubert@deploy2003: helmfile [codfw] DONE helmfile.d/services/mobileapps: apply
* 12:44 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 12:44 cgoubert@deploy2003: helmfile [codfw] START helmfile.d/services/mobileapps: apply
* 12:43 cgoubert@deploy2003: helmfile [staging] DONE helmfile.d/services/mobileapps: apply
* 12:43 cgoubert@deploy2003: helmfile [staging] START helmfile.d/services/mobileapps: apply
* 12:42 kevinbazira@deploy2003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:24 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast4006.wikimedia.org
* 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt2002.wikimedia.org
* 12:20 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host bast7002.wikimedia.org
* 12:18 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast4006.wikimedia.org
* 12:18 bwojtowicz@deploy2003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 12:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1004
* 12:18 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1004
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt2002.wikimedia.org
* 12:14 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host bast7002.wikimedia.org
* 12:10 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[2003,2014].codfw.wmnet with reason: reboot & upgrade
* 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host apt-staging2001.codfw.wmnet
* 12:09 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid1003.eqiad.wmnet
* 12:06 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host apt-staging2001.codfw.wmnet
* 12:05 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid1003.eqiad.wmnet
* 12:03 jynus@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 6:00:00 on backup[1003,1014].eqiad.wmnet with reason: reboot & upgrade
* 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host failoid2003.codfw.wmnet
* 12:02 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc1003.wikimedia.org
* 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host failoid2003.codfw.wmnet
* 11:58 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc1003.wikimedia.org
* 11:55 jmm@dns1004: END - running authdns-update
* 11:53 jmm@dns1004: START - running authdns-update
* 11:50 jmm@dns1004: END - running authdns-update
* 11:48 jmm@dns1004: START - running authdns-update
* 11:27 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host irc2003.wikimedia.org
* 11:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host irc2003.wikimedia.org
* 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint2001.codfw.wmnet
* 11:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-maint1001.eqiad.wmnet
* 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint2001.codfw.wmnet
* 11:12 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-maint1001.eqiad.wmnet
* 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw2001.wikimedia.org
* 11:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host ldap-rw1001.wikimedia.org
* 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw2001.wikimedia.org
* 11:04 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host ldap-rw1001.wikimedia.org
* 11:03 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon1003.wikimedia.org
* 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2005.codfw.wmnet
* 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:00 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 10:59 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2005.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 10:57 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon1003.wikimedia.org
* 10:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netmon2002.wikimedia.org
* 10:55 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 10:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netmon2002.wikimedia.org
* 10:51 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:50 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2005.codfw.wmnet
* 10:41 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab2002.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:40 cmooney@cumin1003: END (PASS) - Cookbook sre.network.host-bgp (exit_code=0) for host ml-serve1003
* 10:40 cmooney@cumin1003: START - Cookbook sre.network.host-bgp for host ml-serve1003
* 10:39 aokoth@cumin1003: END (PASS) - Cookbook sre.gitlab.upgrade (exit_code=0) on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:38 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2033.codfw.wmnet
* 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install2005.wikimedia.org
* 10:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install1005.wikimedia.org
* 10:35 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1004.eqiad.wmnet with OS bookworm
* 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install1005.wikimedia.org
* 10:31 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install2005.wikimedia.org
* 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install4004.wikimedia.org
* 10:30 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install3004.wikimedia.org
* 10:29 aokoth@cumin1003: START - Cookbook sre.gitlab.upgrade on GitLab host gitlab1003.wikimedia.org with reason: Security Release - [[phab:T431656|T431656]]
* 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install3004.wikimedia.org
* 10:24 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install4004.wikimedia.org
* 10:23 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2033.codfw.wmnet
* 10:21 moritzm: failover Ganeti master in codfw/routed to ganeti2034 [[phab:T430928|T430928]]
* 10:19 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.addnode (exit_code=0) for new host ganeti2031.codfw.wmnet to cluster codfw and group B
* 10:19 moritzm: readded ganeti2031 to the codfw Ganeti cluster [[phab:T430910|T430910]]
* 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage
* 10:18 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install5004.wikimedia.org
* 10:18 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003
* 10:18 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003
* 10:17 jmm@cumin2003: START - Cookbook sre.ganeti.addnode for new host ganeti2031.codfw.wmnet to cluster codfw and group B
* 10:17 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install6003.wikimedia.org
* 10:16 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install5004.wikimedia.org
* 10:15 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install6003.wikimedia.org
* 10:15 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1004.eqiad.wmnet with reason: host reimage
* 10:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet
* 10:14 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2008.wikimedia.org
* 10:00 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1004
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1004
* 09:57 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 09:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host install7002.wikimedia.org
* 09:57 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1004
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:57 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1004.eqiad.wmnet 50.48.64.10.in-addr.arpa 0.5.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:57 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003"
* 09:56 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1004 - klausman@cumin1003"
* 09:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet
* 09:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 09:55 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host netflow7002.magru.wmnet
* 09:52 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-coord1001.eqiad.wmnet
* 09:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 09:52 klausman@cumin1003: START - Cookbook sre.dns.netbox
* 09:50 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host install7002.wikimedia.org
* 09:50 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1004
* 09:50 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1004.eqiad.wmnet with OS bookworm
* 09:50 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm
* 09:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host an-test-coord1001.eqiad.wmnet
* 09:49 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-coord1001.eqiad.wmnet
* 09:49 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2008.wikimedia.org
* 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts testvm2007.codfw.wmnet
* 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:49 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:49 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host netflow7002.magru.wmnet
* 09:49 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: testvm2007.codfw.wmnet decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 09:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard1003.eqiad.wmnet
* 09:39 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetboard2003.codfw.wmnet
* 09:39 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 09:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard1003.eqiad.wmnet
* 09:37 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor1003.eqiad.wmnet
* 09:35 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetboard2003.codfw.wmnet
* 09:34 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts testvm2007.codfw.wmnet
* 09:33 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor1003.eqiad.wmnet
* 09:33 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor-dev2001.codfw.wmnet
* 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host debmonitor2003.codfw.wmnet
* 09:29 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host sretest1006.eqiad.wmnet
* 09:27 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor-dev2001.codfw.wmnet
* 09:25 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host debmonitor2003.codfw.wmnet
* 09:23 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] (duration: 06m 27s)
* 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2205: codfw rack B4 repool after maintenance
* 09:23 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host sretest1006.eqiad.wmnet
* 09:23 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2204: codfw rack B4 repool after maintenance
* 09:19 urbanecm@deploy2003: urbanecm: Continuing with deployment
* 09:19 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:18 jmm@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 6 hosts with reason: reboot
* 09:17 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1309102{{!}}NewcomerTasks: Don't fatal on an unconfigured conversion-map fallback (T431668)]]
* 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host ml-serve1003
* 09:08 klausman@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host ml-serve1003
* 09:07 klausman@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host ml-serve1003
* 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:07 klausman@cumin1003: START - Cookbook sre.dns.wipe-cache ml-serve1003.eqiad.wmnet 81.32.64.10.in-addr.arpa 1.8.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 09:07 klausman@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003"
* 09:06 klausman@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host ml-serve1003 - klausman@cumin1003"
* 08:58 klausman@cumin1003: START - Cookbook sre.dns.netbox
* 08:57 klausman@cumin1003: START - Cookbook sre.hosts.move-vlan for host ml-serve1003
* 08:57 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm
* 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.roll-reimage-nodes (exit_code=0) rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve1003.eqiad.wmnet
* 08:55 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve1003.eqiad.wmnet
* 08:55 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host ml-serve1003.eqiad.wmnet with OS bookworm
* 08:39 klausman@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 08:38 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2205: codfw rack B4 repool after maintenance
* 08:37 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool db2204: codfw rack B4 repool after maintenance
* 08:36 klausman@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on ml-serve1003.eqiad.wmnet with reason: host reimage
* 08:35 hashar@deploy2003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 08:32 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:32 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:31 hashar@deploy2003: Rolling back deployment
* 08:26 moritzm: failover Ganeti master in codfw to ganeti2048 [[phab:T430928|T430928]]
* 08:16 klausman@cumin1003: START - Cookbook sre.hosts.reimage for host ml-serve1003.eqiad.wmnet with OS bookworm
* 08:16 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2004.codfw.wmnet
* 08:16 klausman@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve1003.eqiad.wmnet
* 08:16 klausman@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve1003.eqiad.wmnet
* 08:16 klausman@cumin1003: START - Cookbook sre.k8s.roll-reimage-nodes rolling reimage on P<nowiki>{</nowiki>ml-serve1003.eqiad.wmnet<nowiki>}</nowiki> and (A:ml-serve-master-eqiad or A:ml-serve-worker-eqiad)
* 08:15 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host build2002.codfw.wmnet
* 08:15 XioNoX: lsw1-b4-codfw> request system reboot - [[phab:T430910|T430910]]
* 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b4-codfw,lsw1-b4-codfw IPv6,lsw1-b4-codfw.mgmt with reason: Switch maintenance
* 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B4
* 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:10 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2004.codfw.wmnet
* 08:10 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2129,2137-2138,2156,2270-2271].codfw.wmnet
* 08:09 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host build2002.codfw.wmnet
* 08:09 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2205: codfw rack B4 depool for maintenance
* 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2205: codfw rack B4 depool for maintenance
* 08:08 jmm@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2003.codfw.wmnet
* 08:08 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2204: codfw rack B4 depool for maintenance
* 08:08 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool db2204: codfw rack B4 depool for maintenance
* 08:08 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on 27 hosts with reason: codfw rack B4 depool for maintenance
* 08:03 jmm@cumin2002: START - Cookbook sre.hosts.reboot-single for host cumin2003.codfw.wmnet
* 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B4
* 07:54 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1008.eqiad.wmnet with OS trixie
* 07:49 wmde-fisch@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] (duration: 08m 36s)
* 07:44 wmde-fisch@deploy2003: wmde-fisch: Continuing with deployment
* 07:43 wmde-fisch@deploy2003: wmde-fisch: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:41 wmde-fisch@deploy2003: Started scap sync-world: Backport for [[gerrit:1308687{{!}}Enable sub-references on more group2 wikis (batch2) (T430941)]]
* 07:35 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage
* 07:31 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1008.eqiad.wmnet with reason: host reimage
* 07:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie
* 07:10 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 07:00 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1008.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 06:59 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1008.eqiad.wmnet with OS trixie
* 06:57 Emperor: rebalance thanos swift rings after previous re-image of thanos-fe1004 to trixie
* 06:47 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1008.eqiad.wmnet with OS trixie
* 04:10 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 14 days, 0:00:00 on cp6008.drmrs.wmnet with reason: Hardware failure - [[phab:T431651|T431651]]
* 03:55 brett@puppetserver1001: conftool action : set/pooled=no; selector: name=cp6008.*
* 03:29 ryankemper: [[phab:T431311|T431311]] Repooled eqiad cirrussearch clusters (`chi/omega/psi`) following completion of OpenSearch 2.19 migration
* 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=eqiad
* 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=eqiad
* 03:27 ryankemper@cumin2002: conftool action : set/pooled=true; selector: dnsdisc=search,name=eqiad
* 02:07 mwpresync@deploy2003: Finished scap build-images: Publishing wmf/next image (duration: 06m 31s)
* 02:00 mwpresync@deploy2003: Started scap build-images: Publishing wmf/next image
== 2026-07-08 ==
* 23:52 Amir1: ladsgroup@deploy2003:~$ mwscript-k8s --follow -- extensions/ORES/maintenance/PurgeScoreCache.php --wiki=simplewiki --model damaging --old ([[phab:T431159|T431159]])
* 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 23:46 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003"
* 23:46 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add missing PTR for 2001:df2:e500:fe08::1 - cmooney@cumin1003"
* 23:40 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 23:16 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 23:15 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 22:42 rzl@deploy2003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 22:40 urbanecm@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] (duration: 12m 55s)
* 22:40 rzl@deploy2003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 22:37 rzl@deploy2003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 22:36 rzl@deploy2003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 22:35 rzl@deploy2003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 22:34 urbanecm@deploy2003: urbanecm: Continuing with deployment
* 22:33 urbanecm@deploy2003: urbanecm: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 22:33 rzl@deploy2003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 22:32 rzl@deploy2003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 22:30 rzl@deploy2003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 22:30 rzl@deploy2003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 22:29 rzl@deploy2003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 22:27 urbanecm@deploy2003: Started scap sync-world: Backport for [[gerrit:1308782{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]], [[gerrit:1308781{{!}}fix: Return consistent tuple from filterLimitReachedAddLink (T431625)]]
* 22:26 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 22:22 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 22:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 22:19 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 22:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 22:17 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 22:14 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 22:13 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2002.codfw.wmnet with reason: host reimage
* 22:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 22:13 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 22:09 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 22:06 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 22:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1094.eqiad.wmnet with OS trixie
* 22:01 urbanecm: Make https://test.wikipedia.org/w/index.php?title=MediaWiki:GrowthExperimentsSuggestedEdits.json&diff=prev&oldid=750552 with GrowthExperiments disabled (via mw-experimental), then run `\MediaWiki\MediaWikiServices::getInstance()->get('CommunityConfiguration.ProviderFactory')->newProvider('GrowthSuggestedEdits')->getStore()->invalidate()` ([[phab:T431625|T431625]])
* 21:56 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d2-codfw
* 21:55 urbanecm@deploy2003: helmfile [eqiad] DONE helmfile.d/services/mw-experimental: apply
* 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d2-codfw
* 21:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c4-codfw
* 21:55 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c4-codfw
* 21:55 urbanecm@deploy2003: helmfile [eqiad] START helmfile.d/services/mw-experimental: apply
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2002
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2002
* 21:54 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2002
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:54 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2002.codfw.wmnet 50.32.192.10.in-addr.arpa 0.5.0.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:54 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003"
* 21:54 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2002 - bking@cumin2003"
* 21:49 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:49 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2002
* 21:49 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2002.codfw.wmnet with OS trixie
* 21:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage
* 21:42 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 21:39 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 21:37 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1094.eqiad.wmnet with reason: host reimage
* 21:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 21:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 21:29 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 21:27 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 21:22 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1094.eqiad.wmnet with OS trixie
* 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host restbase2039.codfw.wmnet with OS bullseye
* 21:21 jhancock@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002"
* 21:21 jhancock@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jhancock@cumin2002"
* 21:04 jhancock@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage
* 21:00 jhancock@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on restbase2039.codfw.wmnet with reason: host reimage
* 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1073.eqiad.wmnet with OS trixie
* 20:48 mutante: deploy2003 - kill 1102 (stunnel4) ; systemctl start stunnel4 ([[phab:T418262|T418262]])
* 20:42 cjming@deploy2003: Finished scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] (duration: 33m 02s)
* 20:42 jhancock@cumin2002: START - Cookbook sre.hosts.reimage for host restbase2039.codfw.wmnet with OS bullseye
* 20:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage
* 20:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1073.eqiad.wmnet with reason: host reimage
* 20:30 cjming@deploy2003: cjming: Continuing with deployment
* 20:28 cjming@deploy2003: cjming: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1098.eqiad.wmnet with OS trixie
* 20:13 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1073.eqiad.wmnet with OS trixie
* 20:09 cjming@deploy2003: Started scap sync-world: Backport for [[gerrit:1308183{{!}}Move Test Kitchen config from CommonSettings.php (T431257)]]
* 20:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage
* 19:56 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1098.eqiad.wmnet with reason: host reimage
* 19:55 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d4-codfw
* 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d4-codfw
* 19:54 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c1-codfw
* 19:54 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c1-codfw
* 19:52 mutante: restarting gerrit on gerrit.wikimedia.org (gerrit2003)
* 19:48 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker2331.codfw.wmnet
* 19:48 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker2331.codfw.wmnet
* 19:48 mutante: restarting gerrit on gerrit-replica.wikimedia.org (gerrit1003)
* 19:46 mutante: restarting gerrit on gerrit-spare.wikimedia.org (gerrit2002)
* 19:43 jasmine@cumin2002: conftool action : set/pooled=yes; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc
* 19:43 jasmine@cumin2002: conftool action : set/weight=10; selector: name=wikikube-worker2331.codfw.wmnet,cluster=kubernetes,service=kubesvc
* 19:40 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1098.eqiad.wmnet with OS trixie
* 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d5-codfw
* 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d5-codfw
* 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c7-codfw
* 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c7-codfw
* 19:38 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c5-codfw
* 19:38 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c5-codfw
* 19:30 jasmine_: ran homer on lsw1-d8-codfw, adding wikikube-worker2331 to cluster
* 19:29 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1100.eqiad.wmnet with OS trixie
* 19:20 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d8-codfw
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d8-codfw
* 19:19 mutante: gerrit - replacing private key for registerEmail verification
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr2-magru
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr2-magru
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d7-codfw
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d7-codfw
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d3-codfw
* 19:19 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d3-codfw
* 19:19 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d1-codfw
* 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d1-codfw
* 19:18 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c2-codfw
* 19:18 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c2-codfw
* 19:11 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device ssw1-d1-codfw
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device ssw1-d1-codfw
* 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device cr1-magru
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device cr1-magru
* 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d8-codfw
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d8-codfw
* 19:10 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-d6-codfw
* 19:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage
* 19:10 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-d6-codfw
* 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c6-codfw
* 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c6-codfw
* 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device lsw1-c3-codfw
* 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device lsw1-c3-codfw
* 19:09 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b4-magru
* 19:09 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b4-magru
* 19:08 cmooney@cumin1003: END (PASS) - Cookbook sre.network.tls (exit_code=0) for network device asw1-b3-magru
* 19:08 cmooney@cumin1003: START - Cookbook sre.network.tls for network device asw1-b3-magru
* 19:05 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1100.eqiad.wmnet with reason: host reimage
* 19:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1122.eqiad.wmnet with OS trixie
* 19:00 rzl@deploy2003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 18:59 topranks: rolling out update to BGP ACL on Nokia Switches eqiad, codfw & ulsfo [[phab:T425703|T425703]]
* 18:58 rzl@deploy2003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 18:57 rzl@deploy2003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 18:55 rzl@deploy2003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 18:53 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:52 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1100.eqiad.wmnet with OS trixie
* 18:48 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1068.eqiad.wmnet with OS trixie
* 18:47 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie
* 18:47 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:46 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage
* 18:38 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1122.eqiad.wmnet with reason: host reimage
* 18:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage
* 18:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1122
* 18:26 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1122
* 18:25 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1122
* 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:25 bking@cumin2003: START - Cookbook sre.dns.wipe-cache cirrussearch1122.eqiad.wmnet 31.48.64.10.in-addr.arpa 1.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 18:25 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 18:25 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003"
* 18:25 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1122 - bking@cumin2003"
* 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 18:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1068.eqiad.wmnet with reason: host reimage
* 18:21 rzl@deploy2003: helmfile [aux-k8s-codfw] START helmfile.d/admin 'apply'.
* 18:21 rzl@deploy2003: helmfile [aux-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 18:20 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 18:19 rzl@deploy2003: helmfile [aux-k8s-eqiad] START helmfile.d/admin 'apply'.
* 18:19 bking@cumin2003: START - Cookbook sre.dns.netbox
* 18:18 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1122
* 18:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1122.eqiad.wmnet with OS trixie
* 18:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 18:15 rzl@deploy2003: helmfile [staging-codfw] DONE helmfile.d/admin 'apply'.
* 18:13 rzl@deploy2003: helmfile [staging-codfw] START helmfile.d/admin 'apply'.
* 18:13 rzl@deploy2003: helmfile [staging-eqiad] DONE helmfile.d/admin 'apply'.
* 18:10 rzl@deploy2003: helmfile [staging-eqiad] START helmfile.d/admin 'apply'.
* 18:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1068.eqiad.wmnet with OS trixie
* 18:01 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 18m 29s)
* 18:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:55 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:42 kamila@deploy2003: Started scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]]
* 17:42 kamila@deploy2003: Finished scap sync-world: Test deployment to validate deployment server switchover - [[phab:T423714|T423714]] (duration: 19m 50s)
* 17:42 kamila@deploy2003: Rolling back deployment
* 17:35 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1102.eqiad.wmnet with OS trixie
* 17:31 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 17:18 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 17:16 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 22m 07s)
* 17:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 17:11 kamila@dns1005: END - running authdns-update
* 17:09 kamila@dns1005: START - running authdns-update
* 17:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1102.eqiad.wmnet with reason: host reimage
* 17:04 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1164.eqiad.wmnet
* 17:04 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1164.eqiad.wmnet
* 17:04 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1164.eqiad.wmnet
* 16:56 kamila@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on releases2003.codfw.wmnet,releases1003.eqiad.wmnet with reason: Deployment server switchover
* 16:54 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server
* 16:53 kamila@deploy1003: Unlocked for deployment [MediaWiki]: switching deployment server (duration: 04m 02s)
* 16:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1102.eqiad.wmnet with OS trixie
* 16:49 kamila@deploy1003: Locking from deployment [MediaWiki]: switching deployment server
* 16:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1095.eqiad.wmnet with OS trixie
* 16:45 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1093.eqiad.wmnet with OS trixie
* 16:43 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1164.eqiad.wmnet with OS trixie
* 16:27 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage
* 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 16:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 16:23 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage
* 16:18 bking@cumin2003: END (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage
* 16:16 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1164.eqiad.wmnet with reason: host reimage
* 16:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1095.eqiad.wmnet with reason: host reimage
* 16:09 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 16:09 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 16:08 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1093.eqiad.wmnet with reason: host reimage
* 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Pool test
* 15:59 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Pool test
* 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Depool test
* 15:58 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:58 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Depool test
* 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1164
* 15:57 jasmine@cumin2002: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1164
* 15:57 jasmine@cumin2002: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1164
* 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:56 jasmine@cumin2002: START - Cookbook sre.dns.wipe-cache wikikube-worker1164.eqiad.wmnet 114.48.64.10.in-addr.arpa 4.1.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:56 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002"
* 15:56 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1164 - jasmine@cumin2002"
* 15:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1093.eqiad.wmnet with OS trixie
* 15:51 jasmine@cumin2002: START - Cookbook sre.dns.netbox
* 15:51 jasmine@cumin2002: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1164
* 15:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-worker1164.eqiad.wmnet with OS trixie
* 15:50 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet
* 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet
* 15:50 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet
* 15:48 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1095.eqiad.wmnet with OS trixie
* 15:42 jasmine@cumin2002: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=99) Renumbering for host wikikube-worker1164.eqiad.wmnet
* 15:42 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1164.eqiad.wmnet
* 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:42 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1164.eqiad.wmnet
* 15:42 jasmine@cumin2002: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1164.eqiad.wmnet
* 15:39 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:39 elukey@cumin1003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 15:36 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1007.eqiad.wmnet with OS trixie
* 15:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:19 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage
* 15:15 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-pretrain: apply
* 15:15 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet
* 15:15 blake@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-pretrain: apply
* 15:15 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:15 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1007.eqiad.wmnet with reason: host reimage
* 15:15 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test
* 15:14 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet
* 15:14 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2228: Depool test
* 15:14 fceratto@cumin1003: START - Cookbook sre.mysql.depool depool db2228: Depool test
* 15:10 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 15:08 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 15:08 blake@deploy1003: helmfile [eqiad] DONE helmfile.d/admin 'apply'.
* 15:06 blake@deploy1003: helmfile [eqiad] START helmfile.d/admin 'apply'.
* 15:06 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet
* 15:06 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 15:06 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:05 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet
* 15:05 sukhe@cumin1003: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 15:04 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:04 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:04 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:03 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 15:03 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 15:03 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:03 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]]
* 15:03 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 15:03 sukhe@cumin1003: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 15:03 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]]
* 15:03 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1002.eqiad.wmnet
* 15:02 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet
* 14:59 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:59 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host an-test-client1002.eqiad.wmnet,an-test-coord[1001-1002].eqiad.wmnet,an-test-druid1001.eqiad.wmnet,an-test-master[1001-1004].eqiad.wmnet,an-test-presto1001.eqiad.wmnet,an-test-ui1001.eqiad.wmnet,an-test-worker[1001-1003].eqiad.wmnet
* 14:55 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie
* 14:52 swfrench-wmf: restarted ulsfo confds, confirmed now connected to codfw backends except those using wikimedia.org SRV record - [[phab:T430909|T430909]]
* 14:51 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:50 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:49 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:47 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:46 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:44 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:43 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:41 moritzm: uninstalling dhcpcd-base from trixie hosts which still have it installed [[phab:T414341|T414341]]
* 14:40 sukhe: sudo cumin -b1 -s120 "P<nowiki>{</nowiki>lvs2011*<nowiki>}</nowiki> or P<nowiki>{</nowiki>lvs2012*<nowiki>}</nowiki>" "systemctl restart pybal.service"
* 14:39 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1007.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 14:39 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1007.eqiad.wmnet with OS trixie
* 14:37 sukhe: restart pybal on lvs2013 to revert back to conf2004
* 14:35 sukhe: restart pybal on lvs2014 to revert back to conf2004
* 14:34 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records back to codfw - [[phab:T430909|T430909]]
* 14:32 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet
* 14:32 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet
* 14:31 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1007.eqiad.wmnet with OS trixie
* 14:31 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:31 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:31 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:31 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:30 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:30 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:30 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2228: Pool test
* 14:30 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:29 swfrench@dns1004: END - running authdns-update
* 14:29 moritzm: installing jackson-core security updates
* 14:27 swfrench@dns1004: START - running authdns-update
* 14:22 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:22 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1119.eqiad.wmnet with OS trixie
* 14:22 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:21 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:20 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:20 blake@deploy1003: helmfile [codfw] DONE helmfile.d/admin 'apply'.
* 14:20 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:19 blake@deploy1003: helmfile [codfw] START helmfile.d/admin 'apply'.
* 14:19 moritzm: installing librabbitmq security updates
* 14:19 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1002.eqiad.wmnet
* 14:18 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1002.eqiad.wmnet
* 14:16 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:16 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:16 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:15 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:15 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:14 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2228: Pool test
* 14:14 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:14 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[1002-1003].eqiad.wmnet
* 14:10 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 14:08 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[1002-1003].eqiad.wmnet
* 14:05 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1118.eqiad.wmnet with OS trixie
* 14:02 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test1001.eqiad.wmnet
* 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@eqiad
* 14:00 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:59 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 13:58 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 13:57 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage
* 13:54 moritzm: installing libcap2 security updates
* 13:53 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1119.eqiad.wmnet with reason: host reimage
* 13:52 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet
* 13:52 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0)
* 13:52 fceratto@cumin1003: START - Cookbook sre.mysql.depool
* 13:50 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@eqiad
* 13:50 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1051.eqiad.wmnet
* 13:50 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1051.eqiad.wmnet
* 13:50 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1051.eqiad.wmnet
* 13:49 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 13:45 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1081.eqiad.wmnet with OS trixie
* 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1119
* 13:41 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1119
* 13:40 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1119
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1119.eqiad.wmnet 97.32.64.10.in-addr.arpa 7.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003"
* 13:40 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1119 - atsuko@cumin1003"
* 13:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage
* 13:39 moritzm: installing krb5 security updates
* 13:37 Lucas_WMDE: UTC afternoon backport+config window done
* 13:37 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1006.eqiad.wmnet with OS trixie
* 13:36 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1118.eqiad.wmnet with reason: host reimage
* 13:36 atsuko@cumin1003: START - Cookbook sre.dns.netbox
* 13:35 sbisson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] (duration: 07m 46s)
* 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1119
* 13:34 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1119.eqiad.wmnet with OS trixie
* 13:30 sbisson@deploy1003: sbisson: Continuing with deployment
* 13:30 moritzm: installing openssh security updates
* 13:30 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough
* 13:29 sbisson@deploy1003: sbisson: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 13:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 13:27 sbisson@deploy1003: Started scap sync-world: Backport for [[gerrit:1308622{{!}}Enable Article Guidance extension on itwiki (T431540)]]
* 13:26 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1051.eqiad.wmnet with OS trixie
* 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage
* 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1118
* 13:24 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1118
* 13:22 stran@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] (duration: 12m 12s)
* 13:21 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1081.eqiad.wmnet with reason: host reimage
* 13:18 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage
* 13:18 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1118
* 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:18 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1118.eqiad.wmnet 90.32.64.10.in-addr.arpa 0.9.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 13:18 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003"
* 13:18 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1118 - atsuko@cumin1003"
* 13:17 stran@deploy1003: stran: Continuing with deployment
* 13:16 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough
* 13:15 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:13 atsuko@cumin1003: START - Cookbook sre.dns.netbox
* 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1118
* 13:12 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1006.eqiad.wmnet with reason: host reimage
* 13:12 stran@deploy1003: stran: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:12 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1118.eqiad.wmnet with OS trixie
* 13:10 stran@deploy1003: Started scap sync-world: Backport for [[gerrit:1308553{{!}}Deploy IRS to enwiki (T431316)]]
* 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-worker@codfw
* 13:05 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:05 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1081.eqiad.wmnet with OS trixie
* 13:05 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage
* 13:04 moritzm: installing jq security updates
* 13:04 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 13:01 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1051.eqiad.wmnet with reason: host reimage
* 12:58 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-worker@codfw
* 12:52 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:50 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1051
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1051
* 12:43 moritzm: installing Python 3.11 security updates
* 12:43 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1051
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:43 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1051.eqiad.wmnet 46.32.64.10.in-addr.arpa 6.4.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:43 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003"
* 12:43 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1051 - blake@cumin1003"
* 12:38 blake@cumin1003: START - Cookbook sre.dns.netbox
* 12:38 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1051
* 12:38 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1051.eqiad.wmnet with OS trixie
* 12:37 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1051.eqiad.wmnet
* 12:36 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1051.eqiad.wmnet
* 12:36 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1051.eqiad.wmnet
* 12:34 mvernon@cumin2003: START - Cookbook sre.hosts.provision for host thanos-be1006.mgmt.eqiad.wmnet with chassis set policy GRACEFUL_RESTART and with Dell SCP reboot policy GRACEFUL
* 12:34 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:27 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:27 mvernon@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:07 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:02 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 12:01 mvernon@cumin2003: END (FAIL) - Cookbook sre.hosts.reimage (exit_code=99) for host thanos-be1006.eqiad.wmnet with OS trixie
* 11:43 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1006.eqiad.wmnet with OS trixie
* 11:38 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1076.eqiad.wmnet with OS trixie
* 11:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1075.eqiad.wmnet with OS trixie
* 11:21 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2047.codfw.wmnet
* 11:19 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2047.codfw.wmnet
* 11:19 moritzm: temporarily remove ganeti2031 from codfw cluster [[phab:T430910|T430910]]
* 11:08 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage
* 11:08 moritzm: installing Linux 6.1.176 on Bookworm servers
* 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1076.eqiad.wmnet with reason: host reimage
* 11:00 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage
* 10:56 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1075.eqiad.wmnet with reason: host reimage
* 10:47 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1076.eqiad.wmnet with OS trixie
* 10:46 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1074.eqiad.wmnet with OS trixie
* 10:45 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-be1005.eqiad.wmnet with OS trixie
* 10:40 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1075.eqiad.wmnet with OS trixie
* 10:32 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet
* 10:29 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage
* 10:25 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-be1005.eqiad.wmnet with reason: host reimage
* 10:24 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage
* 10:17 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1074.eqiad.wmnet with reason: host reimage
* 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-be1005.eqiad.wmnet with OS trixie
* 10:12 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet
* 10:04 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 10:01 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test1001.eqiad.wmnet
* 10:01 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1074.eqiad.wmnet with OS trixie
* 10:01 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test1001.eqiad.wmnet
* 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 09:43 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply
* 09:43 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 09:35 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply
* 09:34 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 09:34 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 09:33 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Test
* 09:32 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 09:32 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 09:31 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool es2051: codfw rack B3 pool after maintenance
* 09:27 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 09:26 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 09:07 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 09:07 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 09:02 ladsgroup@cumin1003: END (PASS) - Cookbook sre.mysql.sanitarium_restart (exit_code=0)
* 08:57 topranks: merge patch to shift eqiad <-> esams traffic onto new 40G circuit
* 08:54 hashar@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 08:50 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1004.eqiad.wmnet with OS trixie
* 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart
* 08:50 ladsgroup@cumin1003: END (FAIL) - Cookbook sre.mysql.sanitarium_restart (exit_code=99)
* 08:50 ladsgroup@cumin1003: START - Cookbook sre.mysql.sanitarium_restart
* 08:45 ayounsi@cumin1003: START - Cookbook sre.mysql.pool pool es2051: codfw rack B3 pool after maintenance
* 08:44 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:44 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:43 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host aux-k8s-worker2007.codfw.wmnet
* 08:43 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host aux-k8s-worker2007.codfw.wmnet
* 08:42 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet
* 08:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2031.codfw.wmnet
* 08:40 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2031.codfw.wmnet
* 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:38 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:35 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.sanitize-wiki (exit_code=97) Managing sanitization for wikis minwikiquote in section s3
* 08:33 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s3
* 08:32 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5
* 08:30 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5
* 08:30 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Managing sanitization for wikis minwikiquote in section s5
* 08:27 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage
* 08:23 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1004.eqiad.wmnet with reason: host reimage
* 08:23 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:22 fceratto@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Managing sanitization for wikis minwikiquote in section s5
* 08:19 XioNoX: lsw1-b3-codfw> request system reboot - [[phab:T430909|T430909]]
* 08:19 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.sanitize-wiki (exit_code=0) Checking sanitization for wikis minwikiquote in section s5
* 08:17 hashar@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 08:16 cwilliams@cumin1003: START - Cookbook sre.mysql.sanitize-wiki Checking sanitization for wikis minwikiquote in section s5
* 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.network.depool-rack (exit_code=0) with action 'depool' for codfw rack B3
* 08:15 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host aux-k8s-worker2007.codfw.wmnet
* 08:15 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lsw1-b3-codfw,lsw1-b3-codfw IPv6,lsw1-b3-codfw.mgmt with reason: Switch maintenance
* 08:15 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host aux-k8s-worker2007.codfw.wmnet
* 08:13 ayounsi@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:07 ayounsi@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2049,2064-2065,2262-2269].codfw.wmnet
* 08:06 ayounsi@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool es2051: codfw rack B3 depool for maintenance
* 08:05 ayounsi@cumin1003: START - Cookbook sre.mysql.depool depool es2051: codfw rack B3 depool for maintenance
* 08:05 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1004.eqiad.wmnet with OS trixie
* 08:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1005.eqiad.wmnet with OS trixie
* 08:03 ayounsi@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on 21 hosts with reason: codfw rack B3 depool for maintenance
* 07:56 ayounsi@cumin1003: START - Cookbook sre.network.depool-rack with action 'depool' for codfw rack B3
* 07:44 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage
* 07:38 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1005.eqiad.wmnet with reason: host reimage
* 07:36 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm
* 07:29 moritzm: installing gnutls28 security updates
* 07:21 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1005.eqiad.wmnet with OS trixie
* 07:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1125.eqiad.wmnet with OS trixie
* 07:13 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage
* 07:07 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1005.eqiad.wmnet with reason: host reimage
* 06:56 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1005.eqiad.wmnet with OS bookworm
* 06:54 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage
* 06:52 elukey: upgrade all trixie hosts to pywmflib 3.1 - [[phab:T430552|T430552]]
* 06:50 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1125.eqiad.wmnet with reason: host reimage
* 06:40 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 06:40 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 06:38 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1125.eqiad.wmnet with OS trixie
* 05:42 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1107.eqiad.wmnet with OS trixie
* 05:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1124.eqiad.wmnet with OS trixie
* 05:31 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1101.eqiad.wmnet with OS trixie
* 05:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage
* 05:17 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage
* 05:13 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1107.eqiad.wmnet with reason: host reimage
* 05:13 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage
* 05:11 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1124.eqiad.wmnet with reason: host reimage
* 05:10 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1101.eqiad.wmnet with reason: host reimage
* 04:58 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1124.eqiad.wmnet with OS trixie
* 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1107.eqiad.wmnet with OS trixie
* 04:55 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch1101.eqiad.wmnet with OS trixie
* 02:27 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] (duration: 08m 14s)
* 02:22 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 02:21 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 02:19 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308251{{!}}Activate minwikiquote (T429922)]]
* 01:59 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] (duration: 09m 46s)
* 01:55 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 01:51 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 01:49 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1308250{{!}}Init minwikiquote (T429922)]]
* 01:03 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1099.eqiad.wmnet with OS trixie
* 00:57 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1110.eqiad.wmnet with OS trixie
* 00:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage
* 00:41 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1099.eqiad.wmnet with reason: host reimage
* 00:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage
* 00:32 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1110.eqiad.wmnet with reason: host reimage
* 00:26 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1099.eqiad.wmnet with OS trixie
* 00:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1110.eqiad.wmnet with OS trixie
== 2026-07-07 ==
* 22:49 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1097.eqiad.wmnet with OS trixie
* 22:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage
* 22:24 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1097.eqiad.wmnet with reason: host reimage
* 22:14 hashar: Restarting Gerrit on gerrit2002 and gerrit1003 (replicas)
* 22:09 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1097.eqiad.wmnet with OS trixie
* 22:07 hashar: Restarting Gerrit on gerrit2003
* 21:13 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-ml: apply
* 21:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1108.eqiad.wmnet with OS trixie
* 20:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1091.eqiad.wmnet with OS trixie
* 20:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1090.eqiad.wmnet with OS trixie
* 20:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage
* 20:36 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage
* 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage
* 20:33 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1091.eqiad.wmnet with reason: host reimage
* 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1108.eqiad.wmnet with reason: host reimage
* 20:30 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1006.eqiad.wmnet
* 20:30 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1090.eqiad.wmnet with reason: host reimage
* 20:30 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1006.eqiad.wmnet
* 20:27 jasmine_: "homer lsw1-c2-eqiad* commit "Added new stacked control plane wikikube-ctrl1006""
* 20:22 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] (duration: 07m 29s)
* 20:20 jasmine_: "homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1006""
* 20:18 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1091.eqiad.wmnet with OS trixie
* 20:17 arlolra@deploy1003: arlolra: Continuing with deployment
* 20:16 arlolra@deploy1003: arlolra: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:16 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie
* 20:15 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie
* 20:14 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1308196{{!}}Bump Parsoid image limit to 1250 (T430854)]]
* 20:09 cwhite: remove 2026-04 swift log archives from centrallog2002 to free some space
* 20:01 bking@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=93) for host cirrussearch1108.eqiad.wmnet with OS trixie
* 19:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1108.eqiad.wmnet with OS trixie
* 19:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1090.eqiad.wmnet with OS trixie
* 19:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1109.eqiad.wmnet with OS trixie
* 19:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1092.eqiad.wmnet with OS trixie
* 19:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1123.eqiad.wmnet with OS trixie
* 19:22 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage
* 19:19 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*
* 19:17 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage
* 19:17 jasmine@dns1004: END - running authdns-update
* 19:15 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1109.eqiad.wmnet with reason: host reimage
* 19:15 jasmine@dns1004: START - running authdns-update
* 19:15 cdobbins@dns1004: END - running authdns-update
* 19:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage
* 19:13 cdobbins@dns1004: START - running authdns-update
* 19:12 cdobbins@cumin2002: conftool action : set/pooled=yes; selector: name=dns7002.*,service=authdns-update
* 19:11 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1092.eqiad.wmnet with reason: host reimage
* 19:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1123.eqiad.wmnet with reason: host reimage
* 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1123.eqiad.wmnet with OS trixie
* 18:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1109.eqiad.wmnet with OS trixie
* 18:56 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1092.eqiad.wmnet with OS trixie
* 18:52 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 18:49 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-eqsin ([[phab:T430909|T430909]])
* 18:40 swfrench@cumin2002: END (PASS) - Cookbook sre.loadbalancer.upgrade (exit_code=0) restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 18:38 swfrench@cumin2002: START - Cookbook sre.loadbalancer.upgrade restart A:liberica-ulsfo ([[phab:T430909|T430909]])
* 18:11 swfrench-wmf: restarted eqsin, codfw confds - [[phab:T430909|T430909]]
* 18:01 swfrench-wmf: restarted navtiming on webperf2003 - [[phab:T430909|T430909]]
* 17:59 swfrench-wmf: restarted ulsfo confds, confirmed now connected to eqiad backends - [[phab:T430909|T430909]]
* 17:52 sukhe: restart pybal on lvs2011 to switch from conf2004 to conf1008: [[phab:T430909|T430909]]
* 17:51 sukhe: restart pybal on lvs2012 to switch from conf2004 to conf1008 [puppet re-enabled there]: [[phab:T430909|T430909]]
* 17:46 sukhe: restart pybal on lvs2013 to switch from conf2004 to conf1008: [[phab:T430909|T430909]]
* 17:44 swfrench-wmf: switched codfw, eqsin, ulsfo etcd client SRV records to eqiad - [[phab:T430909|T430909]]
* 17:43 swfrench@dns1004: END - running authdns-update
* 17:40 swfrench@dns1004: START - running authdns-update
* 17:40 sukhe: restart pybal on lvs2014 to switch from conf2004 to conf1008: [[phab:T430909|T430909]]
* 17:21 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1003.eqiad.wmnet
* 17:15 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1003.eqiad.wmnet
* 17:14 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1002.eqiad.wmnet
* 17:06 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1002.eqiad.wmnet
* 17:06 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-codfw' 'systemctl restart pybal.service' # lvs2013, [[phab:T416623|T416623]]
* 17:04 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cloudrabbit1001.eqiad.wmnet
* 17:02 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1111.eqiad.wmnet with OS trixie
* 17:00 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-codfw' 'systemctl restart pybal.service' # lvs2014, [[phab:T416623|T416623]]
* 16:58 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host cloudrabbit1001.eqiad.wmnet
* 16:58 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host dns7002.wikimedia.org with OS bookworm
* 16:55 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-low-traffic-eqiad' 'systemctl restart pybal.service' # lvs1019, [[phab:T416623|T416623]]
* 16:53 rzl: rzl@cumin2003:~$ sudo cumin 'A:lvs-secondary-eqiad' 'systemctl restart pybal.service' # lvs1020, [[phab:T416623|T416623]]
* 16:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage
* 16:40 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1111.eqiad.wmnet with reason: host reimage
* 16:38 cmooney@cumin1003: END (FAIL) - Cookbook sre.network.peering (exit_code=99) with action 'configure' for AS: 47794
* 16:35 cmooney@cumin1003: START - Cookbook sre.network.peering with action 'configure' for AS: 47794
* 16:27 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1111.eqiad.wmnet with OS trixie
* 16:21 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1006.eqiad.wmnet with OS trixie
* 16:06 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 16:03 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage
* 15:58 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1121.eqiad.wmnet with OS trixie
* 15:58 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 15:56 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1006.eqiad.wmnet with reason: host reimage
* 15:54 mutante: jenkins down in planned maintenance window
* 15:42 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1037.eqiad.wmnet
* 15:42 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1037.eqiad.wmnet
* 15:42 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1037.eqiad.wmnet
* 15:41 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1006.eqiad.wmnet with OS trixie
* 15:34 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage
* 15:33 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS bookworm
* 15:33 cdobbins@cumin1003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host dns7002.wikimedia.org with OS trixie
* 15:30 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1121.eqiad.wmnet with reason: host reimage
* 15:29 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1001.wikimedia.org
* 15:20 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1001.wikimedia.org
* 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host cirrussearch1121
* 15:18 atsuko@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host cirrussearch1121
* 15:18 andrew@cumin2002: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host clouddumps1002.wikimedia.org
* 15:17 atsuko@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host cirrussearch1121
* 15:17 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:17 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply
* 15:16 atsuko@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 atsuko@cumin1003: START - Cookbook sre.dns.wipe-cache cirrussearch1121.eqiad.wmnet 30.48.64.10.in-addr.arpa 0.3.0.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 15:16 atsuko@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003"
* 15:16 atsuko@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host cirrussearch1121 - atsuko@cumin1003"
* 15:14 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1037.eqiad.wmnet with OS trixie
* 15:11 atsuko@cumin1003: START - Cookbook sre.dns.netbox
* 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org
* 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.move-vlan for host cirrussearch1121
* 15:09 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org
* 15:09 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org
* 15:09 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1121.eqiad.wmnet with OS trixie
* 15:08 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe1007.eqiad.wmnet with OS trixie
* 15:08 andrew@cumin2002: END (FAIL) - Cookbook sre.hosts.reboot-single (exit_code=99) for host clouddumps1002.wikimedia.org
* 15:08 andrew@cumin2002: START - Cookbook sre.hosts.reboot-single for host clouddumps1002.wikimedia.org
* 15:05 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]] (duration: 00m 47s)
* 15:04 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab1004 for [[phab:T431440|T431440]]
* 15:03 brennen@deploy1003: Finished deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]] (duration: 00m 51s)
* 15:03 brennen@deploy1003: Started deploy [phabricator/deployment@7e02037]: deploy phab2003 for [[phab:T431440|T431440]]
* 15:00 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f] (duration: 02m 10s)
* 14:58 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (thin): Regular analytics weekly train THIN [analytics/refinery@7d8dc71f]
* 14:58 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f] (duration: 04m 14s)
* 14:54 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage
* 14:53 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71]: Regular analytics weekly train [analytics/refinery@7d8dc71f]
* 14:53 javiermonton@deploy1003: Finished deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f] (duration: 02m 00s)
* 14:51 javiermonton@deploy1003: Started deploy [analytics/refinery@7d8dc71] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@7d8dc71f]
* 14:51 arnaudb@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on phab2003.codfw.wmnet,phab[1004-1006].eqiad.wmnet with reason: maintenance
* 14:51 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1037.eqiad.wmnet with reason: host reimage
* 14:50 JavierMonton: Deploying Refinery at {{Gerrit|7d8dc71f}} for change {{Gerrit|1308087}} / [[phab:T431318|T431318]] - update filerevision table sqoop and table
* 14:49 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage
* 14:42 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe1007.eqiad.wmnet with reason: host reimage
* 14:40 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1083.eqiad.wmnet with OS trixie
* 14:37 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:36 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@eqiad
* 14:35 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 14:34 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1037
* 14:34 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1037
* 14:34 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]]
* 14:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 14:34 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]] (duration: 06m 47s)
* 14:34 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-eqiad or A:lvs-secondary-eqiad) and A:bullseye and A:lvs
* 14:33 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:33 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:32 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1037
* 14:31 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:31 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:29 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:29 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@eqiad
* 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:29 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:28 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:28 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:27 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:27 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308112{{!}}Revert^2 "[Growth] frwiki: Deploy automated mentor list cleaner" (T427386)]]
* 14:26 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.update-replication (exit_code=0)
* 14:26 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe1007.eqiad.wmnet with OS trixie
* 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:26 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.update-replication (exit_code=99)
* 14:26 fceratto@cumin1003: START - Cookbook sre.mysql.update-replication
* 14:25 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:25 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:25 urbanecm@deploy1003: mwscript-k8s job started: GrowthExperiments:cleanMentorList.php --wiki=frwiki # [[phab:T427386|T427386]]
* 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: END (FAIL) - Cookbook sre.dns.wipe-cache (exit_code=99) wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1037.eqiad.wmnet 143.48.64.10.in-addr.arpa 3.4.1.0.8.4.0.0.4.6.0.0.0.1.0.0.7.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 14:24 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:24 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003"
* 14:24 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1037 - blake@cumin1003"
* 14:20 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage
* 14:19 blake@cumin1003: START - Cookbook sre.dns.netbox
* 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.migrate-service-ipip (exit_code=0) for alias: ml-serve-master@codfw
* 14:19 klausman@cumin2002: END (PASS) - Cookbook sre.loadbalancer.restart-pybal (exit_code=0) rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:19 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1037
* 14:18 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1037.eqiad.wmnet with OS trixie
* 14:18 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1037.eqiad.wmnet
* 14:18 klausman@cumin2002: START - Cookbook sre.loadbalancer.restart-pybal rolling-restart of pybal on (A:lvs-low-traffic-codfw or A:lvs-secondary-codfw) and A:bullseye and A:lvs
* 14:18 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1037.eqiad.wmnet
* 14:18 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1037.eqiad.wmnet
* 14:16 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2007.codfw.wmnet with OS trixie
* 14:16 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1083.eqiad.wmnet with reason: host reimage
* 14:15 blake@cumin1003: END (FAIL) - Cookbook sre.k8s.renumber-node (exit_code=1) Renumbering for host wikikube-worker1036.eqiad.wmnet
* 14:15 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker1036.eqiad.wmnet
* 14:14 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker1036.eqiad.wmnet
* 14:12 klausman@cumin2002: START - Cookbook sre.loadbalancer.migrate-service-ipip for alias: ml-serve-master@codfw
* 14:11 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1120.eqiad.wmnet with OS trixie
* 14:05 moritzm: installing distro-info-data updates from trixie/bookworm point releases
* 14:04 fabfur: disable puppet on A:cp-text to selectively apply https://gerrit.wikimedia.org/r/c/operations/puppet/+/1308040
* 14:03 urbanecm@deploy1003: Finished scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] (duration: 27m 48s)
* 14:00 cdobbins@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 14:00 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1083.eqiad.wmnet with OS trixie
* 13:58 urbanecm@deploy1003: urbanecm: Continuing with deployment
* 13:58 urbanecm@deploy1003: urbanecm: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:58 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm
* 13:57 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage
* 13:57 cdobbins@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on dns7002.wikimedia.org with reason: host reimage
* 13:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage
* 13:50 moritzm: installing Linux 5.10.259 on Bullseye hosts
* 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 13:47 cgoubert@deploy1003: helmfile [aux-k8s-eqiad] START helmfile.d/aux-k8s-services/redioscope: apply
* 13:46 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2007.codfw.wmnet with reason: host reimage
* 13:46 cgoubert@deploy1003: helmfile [aux-k8s-codfw] DONE helmfile.d/aux-k8s-services/redioscope: apply
* 13:46 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1120.eqiad.wmnet with reason: host reimage
* 13:46 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:46 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:45 jgiannelos@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 13:44 jgiannelos@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 13:44 cgoubert@deploy1003: helmfile [aux-k8s-codfw] START helmfile.d/aux-k8s-services/redioscope: apply
* 13:40 jgiannelos@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 13:39 jgiannelos@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 13:38 moritzm: installing e2fsprogs updates from Trixie point release
* 13:35 urbanecm@deploy1003: Started scap sync-world: Backport for [[gerrit:1308106{{!}}fix(MentorChangeLogFormatter): Remove unused XSS suppression (T430693)]], [[gerrit:1308105{{!}}fix(MentorListCleaner): Do not access property before inicialization (T430689)]]
* 13:33 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1120.eqiad.wmnet with OS trixie
* 13:33 topranks: reset cr3-eqsin configuration so traffic uses it again after upgrade
* 13:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1088.eqiad.wmnet with OS trixie
* 13:32 cdobbins@cumin1003: START - Cookbook sre.hosts.reimage for host dns7002.wikimedia.org with OS trixie
* 13:32 cdobbins@cumin1003: conftool action : set/pooled=no; selector: name=dns7002.*
* 13:29 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2007.codfw.wmnet with OS trixie
* 13:23 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage
* 13:20 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2006.codfw.wmnet with OS trixie
* 13:18 blake@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-worker1036.eqiad.wmnet with OS trixie
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on aux-k8s-etcd1004.eqiad.wmnet with reason: host reimage
* 13:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 13:16 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 13:15 jayme: Istio is being upgraded from 1.24.2 to 1.29.4 on wikikube staging eqiad and codfw - [[phab:T427401|T427401]]
* 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1087.eqiad.wmnet with OS trixie
* 13:14 topranks: reboot cr3-eqsin to install new JunOS and set PIC 0/0/0 to 100G
* 13:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage
* 13:13 jmm@dns1004: END - running authdns-update
* 13:12 jmm@dns1004: START - running authdns-update
* 13:09 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1088.eqiad.wmnet with reason: host reimage
* 13:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1082.eqiad.wmnet with OS trixie
* 13:07 jmm@dns1004: END - running authdns-update
* 13:06 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1004.eqiad.wmnet with OS bookworm
* 13:05 jmm@dns1004: START - running authdns-update
* 13:02 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage
* 12:58 topranks: load updated JunOS on cr3-eqsin [[phab:T429386|T429386]]
* 12:58 blake@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage
* 12:57 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2001.codfw.wmnet
* 12:57 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2006.codfw.wmnet with reason: host reimage
* 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin
* 12:56 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin
* 12:55 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 12:55 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 12:53 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage
* 12:53 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 12:52 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1088.eqiad.wmnet with OS trixie
* 12:52 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1002.eqiad.wmnet
* 12:52 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:51 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2001.codfw.wmnet
* 12:49 blake@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-worker1036.eqiad.wmnet with reason: host reimage
* 12:48 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 12:48 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1087.eqiad.wmnet with reason: host reimage
* 12:44 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage
* 12:43 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1002.eqiad.wmnet
* 12:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 12:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub: apply<nowiki>}</nowiki>
* 12:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub: apply
* 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:39 filippo@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003"
* 12:39 filippo@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: move dumps-nfs IP to the shared one - filippo@cumin1003"
* 12:39 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2006.codfw.wmnet with OS trixie
* 12:38 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1082.eqiad.wmnet with reason: host reimage
* 12:36 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 12:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 12:33 filippo@cumin1003: START - Cookbook sre.dns.netbox
* 12:32 blake@cumin1003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wikikube-worker1036
* 12:32 blake@cumin1003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wikikube-worker1036
* 12:32 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2005.codfw.wmnet with OS trixie
* 12:32 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1087.eqiad.wmnet with OS trixie
* 12:30 jmm@dns1004: END - running authdns-update
* 12:29 blake@cumin1003: START - Cookbook sre.network.configure-switch-interfaces for host wikikube-worker1036
* 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:29 blake@cumin1003: START - Cookbook sre.dns.wipe-cache wikikube-worker1036.eqiad.wmnet 21.32.64.10.in-addr.arpa 1.2.0.0.2.3.0.0.4.6.0.0.0.1.0.0.3.0.1.0.1.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 12:29 blake@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:29 blake@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003"
* 12:29 blake@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wikikube-worker1036 - blake@cumin1003"
* 12:28 jmm@dns1004: START - running authdns-update
* 12:26 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:26 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:23 blake@cumin1003: START - Cookbook sre.dns.netbox
* 12:23 blake@cumin1003: START - Cookbook sre.hosts.move-vlan for host wikikube-worker1036
* 12:23 blake@cumin1003: START - Cookbook sre.hosts.reimage for host wikikube-worker1036.eqiad.wmnet with OS trixie
* 12:22 blake@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker1036.eqiad.wmnet
* 12:22 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1082.eqiad.wmnet with OS trixie
* 12:22 blake@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker1036.eqiad.wmnet
* 12:22 blake@cumin1003: START - Cookbook sre.k8s.renumber-node Renumbering for host wikikube-worker1036.eqiad.wmnet
* 12:21 marostegui: Restart mariadb@s7 on db1155 to pick up new filters - [[phab:T431124|T431124]]
* 12:21 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: restarting for replication filter
* 12:20 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:19 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:14 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:14 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage
* 12:14 brouberol@cumin1003: END (PASS) - Cookbook sre.wdqs.restart (exit_code=0)
* 12:08 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:07 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:07 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2005.codfw.wmnet with reason: host reimage
* 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:06 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:06 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:05 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:05 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-eqiad
* 12:04 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1002.eqiad.wmnet
* 12:04 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1002.eqiad.wmnet
* 12:04 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:04 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 12:03 brouberol@cumin1003: END (FAIL) - Cookbook sre.wdqs.restart (exit_code=99)
* 12:03 brouberol@cumin1003: START - Cookbook sre.wdqs.restart
* 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1002.eqiad.wmnet
* 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1002.eqiad.wmnet
* 11:59 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:59 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:56 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:56 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl1001.eqiad.wmnet
* 11:56 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-eqiad
* 11:50 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2005.codfw.wmnet with OS trixie
* 11:49 blake@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on wikikube-worker1160.eqiad.wmnet with reason: Verifying matchers for silence
* 11:42 topranks: cr3-eqsin, begin traffic drain to reset PIC and upgrade JunOS [[phab:T429386|T429386]]
* 11:41 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs[5004-5006].eqsin.wmnet with reason: upgrade JunOS cr3-eqsin
* 11:39 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cr1-codfw,cr[2-3]-eqsin,cr3-eqsin IPv6,cr3-eqsin.mgmt with reason: upgrade JunOS cr3-eqsin
* 11:36 mvernon@cumin2003: conftool action : set/pooled=yes; selector: name=thanos-fe2004.codfw.wmnet
* 11:35 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1086.eqiad.wmnet with OS trixie
* 11:35 mvernon@cumin2003: conftool action : set/pooled=inactive; selector: name=thanos-fe2004.codfw.wmnet
* 11:32 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1085.eqiad.wmnet with OS trixie
* 11:14 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage
* 11:10 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage
* 11:04 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host thanos-fe2004.codfw.wmnet with OS trixie
* 11:03 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1086.eqiad.wmnet with reason: host reimage
* 11:02 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1085.eqiad.wmnet with reason: host reimage
* 10:48 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage
* 10:48 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1086.eqiad.wmnet with OS trixie
* 10:46 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1085.eqiad.wmnet with OS trixie
* 10:44 cgoubert@deploy1003: Finished deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 16m 44s)
* 10:43 mvernon@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on thanos-fe2004.codfw.wmnet with reason: host reimage
* 10:35 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:27 cgoubert@deploy1003: Started deploy [restbase/deploy@2fc37d4]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]]
* 10:27 cgoubert@deploy1003: Finished deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]] (duration: 00m 45s)
* 10:26 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1117.eqiad.wmnet with OS trixie
* 10:26 cgoubert@deploy1003: Started deploy [restbase/deploy@8a25036]: {{Gerrit|1306049}}: Add isvwiki to RESTBase {{!}} https://gerrit.wikimedia.org/r/c/mediawiki/services/restbase/deploy/+/1306049 - [[phab:T429936|T429936]]
* 10:26 mvernon@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host thanos-fe2004
* 10:25 mvernon@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host thanos-fe2004
* 10:22 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1116.eqiad.wmnet with OS trixie
* 10:21 mvernon@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host thanos-fe2004
* 10:21 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:20 mvernon@cumin2003: START - Cookbook sre.dns.wipe-cache thanos-fe2004.codfw.wmnet 157.32.192.10.in-addr.arpa 7.5.1.0.2.3.0.0.2.9.1.0.0.1.0.0.3.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:20 mvernon@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003"
* 10:20 mvernon@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host thanos-fe2004 - mvernon@cumin2003"
* 10:15 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2252: Repooling after reboot
* 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.parsercache
* 10:15 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2252: Repooling after reboot
* 10:15 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db1153.eqiad.wmnet
* 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db1153.eqiad.wmnet
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2252.codfw.wmnet
* 10:14 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2252.codfw.wmnet
* 10:12 mvernon@cumin2003: START - Cookbook sre.dns.netbox
* 10:12 mvernon@cumin2003: START - Cookbook sre.hosts.move-vlan for host thanos-fe2004
* 10:11 mvernon@cumin2003: START - Cookbook sre.hosts.reimage for host thanos-fe2004.codfw.wmnet with OS trixie
* 10:07 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage
* 10:03 atsuko@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage
* 09:58 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:58 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1117.eqiad.wmnet with reason: host reimage
* 09:57 atsuko@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1116.eqiad.wmnet with reason: host reimage
* 09:49 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 41 days, 15:00:00 on db2252.codfw.wmnet with reason: Security updates
* 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1117.eqiad.wmnet with OS trixie
* 09:45 atsuko@cumin1003: START - Cookbook sre.hosts.reimage for host cirrussearch1116.eqiad.wmnet with OS trixie
* 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1153: Security updates
* 09:28 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:28 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:28 root@cumin1003: START - Cookbook sre.mysql.depool depool db1153: Security updates
* 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1016: Security updates
* 09:22 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:21 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 09:21 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1016: Security updates
* 09:14 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 09:14 cwilliams@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates
* 08:56 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:56 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:56 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates
* 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:50 javiermonton@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/pageview-trending-relative-next: apply
* 08:45 filippo@dns1006: END - running authdns-update
* 08:43 filippo@dns1006: START - running authdns-update
* 08:42 godog: switch dumps-nfs address to be shared with rsync/http - [[phab:T411248|T411248]]
* 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1016: Security updates
* 08:40 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:40 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:40 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1016: Security updates
* 08:29 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet
* 08:29 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:27 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:27 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:25 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:25 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:17 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host cirrussearch1111.eqiad.wmnet
* 08:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1015: Security updates
* 08:09 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 08:09 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 08:09 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1015: Security updates
* 07:42 Msz2001: Deployed private patch for Suggested Ivestigations
* 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1015: Security updates
* 07:41 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:41 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:41 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1015: Security updates
* 07:40 kevinbazira@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003"
* 07:11 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003
* 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1024: Security updates
* 07:11 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 07:11 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 07:11 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1024: Security updates
* 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fingerprint warnings - oblivian@cumin1003
* 07:10 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fingerprint warnings - oblivian@cumin1003"
* 06:55 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host cirrussearch1111.eqiad.wmnet
* 06:55 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host cirrussearch1111.eqiad.wmnet
* 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1024: Security updates
* 06:48 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 06:48 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 06:48 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates
* 06:42 moritzm: install nginx security updates
* 06:31 root@cumin1003: END (ERROR) - Cookbook sre.mysql.depool (exit_code=97) depool pc1024: Security updates
* 06:21 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1024: Security updates
* 06:19 moritzm: installing php8.2 security updates
* 06:15 moritzm: installing php8.4 security updates
* 04:02 mwpresync@deploy1003: Pruned MediaWiki: 1.47.0-wmf.7 (duration: 02m 38s)
* 03:40 mwpresync@deploy1003: Finished scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]] (duration: 37m 04s)
* 03:03 mwpresync@deploy1003: Started scap sync-world: testwikis to 1.47.0-wmf.10 refs [[phab:T430829|T430829]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 51s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-06 ==
* 23:30 jdlrobson@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] (duration: 09m 39s)
* 23:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1078.eqiad.wmnet with OS trixie
* 23:26 jdlrobson@deploy1003: jdlrobson, bwang: Continuing with deployment
* 23:22 jdlrobson@deploy1003: jdlrobson, bwang: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug)
* 23:21 jdlrobson@deploy1003: Started scap sync-world: Backport for [[gerrit:1306983{{!}}Remove unused user skin preference config (T358273)]], [[gerrit:1306982{{!}}Drop orphaned configuration for Vector skin rollout (T358273)]], [[gerrit:1305921{{!}}Remove wgMinervaEnableSiteNotice config flag (T417638)]], [[gerrit:1306453{{!}}Drop unused VectorNightMode config (T393977)]]
* 23:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage
* 23:06 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1078.eqiad.wmnet with reason: host reimage
* 22:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1078.eqiad.wmnet with OS trixie
* 22:29 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch1114.eqiad.wmnet with reason: reimage on hold until restore completes
* 22:22 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on cirrussearch[1079,1115].eqiad.wmnet with reason: reimage on hold until restore completes
* 21:18 maryum: Deployed security fix for [[phab:T428006|T428006]]
* 20:44 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1079.eqiad.wmnet with OS trixie
* 20:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1077.eqiad.wmnet with OS trixie
* 20:26 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1115.eqiad.wmnet with OS trixie
* 20:25 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage
* 20:21 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1079.eqiad.wmnet with reason: host reimage
* 20:15 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] (duration: 08m 14s)
* 20:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage
* 20:10 krinkle@deploy1003: krinkle, pushpaktiwari: Continuing with deployment
* 20:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage
* 20:08 krinkle@deploy1003: krinkle, pushpaktiwari: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1077.eqiad.wmnet with reason: host reimage
* 20:06 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1303490{{!}}T429269: Send logged-in experiment events to ins-502b]], [[gerrit:1307812{{!}}Re-enable wgTrackMediaRequestProvenance on pilot wikis (group1) (T414338)]]
* 20:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1079.eqiad.wmnet with OS trixie
* 20:04 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1115.eqiad.wmnet with reason: host reimage
* 19:52 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1077.eqiad.wmnet with OS trixie
* 19:51 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1115.eqiad.wmnet with OS trixie
* 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 19:38 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 18:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1114.eqiad.wmnet with OS trixie
* 18:39 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage
* 18:35 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1114.eqiad.wmnet with reason: host reimage
* 18:32 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1112.eqiad.wmnet with OS trixie
* 18:23 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1114.eqiad.wmnet with OS trixie
* 18:21 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1072.eqiad.wmnet with OS trixie
* 18:14 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage
* 18:10 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1112.eqiad.wmnet with reason: host reimage
* 17:59 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage
* 17:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1112.eqiad.wmnet with OS trixie
* 17:55 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1072.eqiad.wmnet with reason: host reimage
* 17:39 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1072.eqiad.wmnet with OS trixie
* 17:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1071.eqiad.wmnet with OS trixie
* 17:18 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1070.eqiad.wmnet with OS trixie
* 17:16 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1084.eqiad.wmnet with OS trixie
* 16:54 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1071.eqiad.wmnet with reason: host reimage
* 16:53 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage
* 16:51 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1070.eqiad.wmnet with reason: host reimage
* 16:49 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1084.eqiad.wmnet with reason: host reimage
* 16:38 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1071.eqiad.wmnet with OS trixie
* 16:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1096.eqiad.wmnet with OS trixie
* 16:35 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1070.eqiad.wmnet with OS trixie
* 16:33 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1084.eqiad.wmnet with OS trixie
* 16:30 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1089.eqiad.wmnet with OS trixie
* 16:28 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1103.eqiad.wmnet with OS trixie
* 16:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage
* 16:14 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1096.eqiad.wmnet with reason: host reimage
* 16:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage
* 16:05 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage
* 16:02 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm
* 16:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1089.eqiad.wmnet with reason: host reimage
* 16:00 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1103.eqiad.wmnet with reason: host reimage
* 15:58 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1096.eqiad.wmnet with OS trixie
* 15:54 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1080.eqiad.wmnet with OS trixie
* 15:46 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1089.eqiad.wmnet with OS trixie
* 15:45 dancy@deploy1003: Installation of scap version "4.272.0" completed for 158 hosts
* 15:43 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1103.eqiad.wmnet with OS trixie
* 15:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1113.eqiad.wmnet with OS trixie
* 15:41 dancy@deploy1003: Installing scap version "4.272.0" for 158 host(s)
* 15:40 klausman@deploy1003: helmfile [ml-serve-eqiad] DONE helmfile.d/admin 'apply'.
* 15:39 klausman@deploy1003: helmfile [ml-serve-eqiad] START helmfile.d/admin 'apply'.
* 15:38 klausman@deploy1003: helmfile [ml-serve-codfw] DONE helmfile.d/admin 'apply'.
* 15:38 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch1069.eqiad.wmnet with OS trixie
* 15:37 klausman@deploy1003: helmfile [ml-serve-codfw] START helmfile.d/admin 'apply'.
* 15:36 klausman@deploy1003: helmfile [ml-staging-codfw] DONE helmfile.d/admin 'apply'.
* 15:34 klausman@deploy1003: helmfile [ml-staging-codfw] START helmfile.d/admin 'apply'.
* 15:33 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage
* 15:27 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1080.eqiad.wmnet with reason: host reimage
* 15:23 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage
* 15:19 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage
* 15:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1113.eqiad.wmnet with reason: host reimage
* 15:16 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch1069.eqiad.wmnet with reason: host reimage
* 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:11 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1080.eqiad.wmnet with OS trixie
* 15:11 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 15:05 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch1113.eqiad.wmnet with OS trixie
* 15:05 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on aux-k8s-etcd1003.eqiad.wmnet with reason: host reimage
* 14:47 elukey@cumin1003: START - Cookbook sre.hosts.reimage for host aux-k8s-etcd1003.eqiad.wmnet with OS bookworm
* 14:33 elukey: rolled out spicerack on all cumin nodes - [[phab:T429699|T429699]]
* 14:32 elukey: upgrade all bookworm hosts to pywmflib 3.1 - [[phab:T430552|T430552]]
* 14:14 marostegui: Setup x4 eqiad topology [[phab:T404715|T404715]]
* 14:13 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 14:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db2230.codfw.wmnet
* 14:07 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db2230.codfw.wmnet
* 13:59 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for db-test[2001-2002].codfw.wmnet
* 13:51 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 13:45 cwilliams@cumin1003: END (ERROR) - Cookbook sre.mysql.major-upgrade (exit_code=97)
* 13:45 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]]
* 13:45 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.reboot-nodes (exit_code=0) rolling reboot on A:ml-serve-master-codfw
* 13:42 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2002.codfw.wmnet
* 13:42 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2002.codfw.wmnet
* 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2002.codfw.wmnet
* 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2002.codfw.wmnet
* 13:38 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host ml-serve-ctrl2001.codfw.wmnet
* 13:38 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host ml-serve-ctrl2001.codfw.wmnet
* 13:35 klausman@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host ml-serve-ctrl2001.codfw.wmnet
* 13:35 klausman@cumin2002: START - Cookbook sre.k8s.pool-depool-node depool for host ml-serve-ctrl2001.codfw.wmnet
* 13:35 klausman@cumin2002: START - Cookbook sre.k8s.reboot-nodes rolling reboot on A:ml-serve-master-codfw
* 12:30 krinkle@deploy1003: Finished scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] (duration: 25m 11s)
* 12:24 krinkle@deploy1003: krinkle: Continuing with deployment
* 12:10 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.drain-node (exit_code=0) for draining ganeti node ganeti2048.codfw.wmnet
* 12:09 krinkle@deploy1003: krinkle: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 12:08 jmm@cumin2003: START - Cookbook sre.ganeti.drain-node for draining ganeti node ganeti2048.codfw.wmnet
* 12:05 krinkle@deploy1003: Started scap sync-world: Backport for [[gerrit:1265672{{!}}robots.php: Change Beta Cluster override from prepend to replace]]
* 11:57 moritzm: installing curl security updates
* 11:49 kevinbazira@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 11:31 moritzm: installing nano security updates
* 11:07 moritzm: failover Ganeti master in codfw to ganeti2032 [[phab:T430909|T430909]]
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 11:04 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host sretest1005.eqiad.wmnet with OS trixie
* 11:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:50 jmm@dns1004: END - running authdns-update
* 10:47 jmm@dns1004: START - running authdns-update
* 10:47 jmm@dns1004: START - running authdns-update
* 10:46 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:44 elukey@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage
* 10:38 elukey@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on sretest1005.eqiad.wmnet with reason: host reimage
* 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:31 marostegui: Setup x4 codfw topology [[phab:T404715|T404715]]
* 10:31 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:24 elukey: spicerack 13.0.0 deployed on cumin2002
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:24 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:23 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 10:21 elukey@cumin2002: START - Cookbook sre.hosts.reimage for host sretest1005.eqiad.wmnet with OS trixie
* 10:20 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:19 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 10:17 elukey: uploaded spicerack_13.0.0 to apt.wikimedia.org bookworm-wikimedia,trixie-wikimedia
* 09:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:52 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:20 elukey: upgrade all bullseye hosts to pywmflib 3.1 - [[phab:T430552|T430552]]
* 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s4
* 09:10 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1015.eqiad.wmnet,service=s6
* 09:07 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 08:58 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 08:56 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/services/mediawiki-dumps-legacy: apply
* 08:28 brouberol@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/services/mediawiki-dumps-legacy: apply
* 08:08 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host cumin2002.codfw.wmnet
* 08:06 godog: remove cloudvirt1046, cloudvirt1062, cloudvirt1074, cloudvirt1075 from maintenance aggregate and put them in network-ovs - [[phab:T424802|T424802]]
* 08:00 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host cumin2002.codfw.wmnet
* 07:58 hashar@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] (duration: 32m 53s)
* 07:57 fabfur: repooled cp4038
* 07:57 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.*
* 07:53 moritzm: installing pyjwt security updates
* 07:47 moritzm: installing openjpeg2 security updates
* 07:45 hashar@deploy1003: vadymts1, hashar: Continuing with deployment
* 07:43 hashar@deploy1003: vadymts1, hashar: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:38 moritzm: installing python-urllib3 security updates
* 07:37 mvernon@cumin2003: END (PASS) - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies (exit_code=0) rolling restart_daemons on A:swift-fe-codfw
* 07:30 fabfur: depooled cp4038 to investigate on possible maxmind failure
* 07:30 fabfur@cumin1003: conftool action : set/pooled=no; selector: name=cp4038.*
* 07:30 fabfur@cumin1003: conftool action : set/pooled=yes; selector: name=cp4038.*
* 07:29 mvernon@cumin2003: START - Cookbook sre.swift.roll-restart-reboot-swift-ms-proxies rolling restart_daemons on A:swift-fe-codfw
* 07:25 hashar@deploy1003: Started scap sync-world: Backport for [[gerrit:1307617{{!}}Turn on PageImages in Author namespace for Ukrainian Wikisource (T431202)]]
* 06:13 moritzm: installing Linux 6.12.95 on trixie hosts
* 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s6
* 05:20 marostegui@cumin1003: conftool action : set/pooled=no; selector: name=clouddb1015.eqiad.wmnet,service=s4
* 05:19 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on clouddb1015.eqiad.wmnet with reason: cloning
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 08s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-05 ==
* 02:01 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 01m 08s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-04 ==
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 58s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-03 ==
* 17:08 topranks: revert protocol preference changes on cr3-ulsfo after upgrade
* 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr2-eqord with reason: upgrade JunOS cr3-ulsfo
* 16:53 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 0:30:00 on cr4-ulsfo with reason: upgrade JunOS cr3-ulsfo
* 16:48 topranks: reboot cr3-ulsfo to upgrade JunOS and reset linecard [[phab:T424839|T424839]]
* 15:52 topranks: adjust outbound BGP policies on cr3-ulsfo to drain router of traffic [[phab:T424839|T424839]]
* 15:45 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on lvs[4008-4010].ulsfo.wmnet with reason: upgrade JunOS cr3-ulsfo
* 15:44 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on asw1-[22-23]-ulsfo,cr3-ulsfo,cr3-ulsfo IPv6,cr3-ulsfo.mgmt with reason: upgrade JunOS cr3-ulsfo
* 15:36 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 15:35 atsuko@cumin1003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 14:40 cmooney@dns3003: END - running authdns-update
* 14:26 cmooney@dns3003: START - running authdns-update
* 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 14:26 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003"
* 14:19 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to ulsfo - cmooney@cumin1003"
* 14:16 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 13:38 sukhe@dns1004: END - running authdns-update
* 13:35 sukhe@dns1004: START - running authdns-update
* 13:26 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 13:26 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 13:26 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:24 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet
* 13:24 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 13:18 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet
* 13:17 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 13:16 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest1005.eqiad.wmnet
* 13:16 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 13:16 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 13:15 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 13:14 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:14 moritzm: imported samplicator 1.3.8rc1-1+deb13u1 to trixie-wikimedia/main [[phab:T337208|T337208]]
* 13:13 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:07 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 13:07 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 13:02 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=99) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:02 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 13:00 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:58 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:57 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 12:57 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 12:53 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:52 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest[2001,2003-2004,2006,2009-2010].codfw.wmnet,sretest[1005-1006].eqiad.wmnet
* 12:50 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.bmc-user-mgmt (exit_code=0) for host sretest1005.eqiad.wmnet
* 12:50 elukey@cumin1003: START - Cookbook sre.hosts.bmc-user-mgmt for host sretest1005.eqiad.wmnet
* 12:47 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 12:41 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 12:40 cgoubert@deploy1003: helmfile [eqiad] DONE helmfile.d/services/ratelimit: apply
* 12:39 cgoubert@deploy1003: helmfile [eqiad] START helmfile.d/services/ratelimit: apply
* 12:32 kamila@cumin1003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host deploy2003.codfw.wmnet
* 12:26 kamila@cumin1003: START - Cookbook sre.hosts.reboot-single for host deploy2003.codfw.wmnet
* 12:23 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org
* 12:19 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org
* 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:15 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[2004-2007].codfw.wmnet
* 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:15 jynus@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003"
* 12:15 jynus@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[2004-2007].codfw.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin2003"
* 12:09 jynus@cumin2003: START - Cookbook sre.dns.netbox
* 11:58 jynus@cumin2003: START - Cookbook sre.hosts.decommission for hosts backup[2004-2007].codfw.wmnet
* 10:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts backup[1004-1007].eqiad.wmnet
* 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:02 jynus@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 10:01 jynus@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: backup[1004-1007].eqiad.wmnet decommissioned, removing all IPs except the asset tag one - jynus@cumin1003"
* 09:52 jynus@cumin1003: START - Cookbook sre.dns.netbox
* 09:39 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:36 jynus@cumin1003: START - Cookbook sre.hosts.decommission for hosts backup[1004-1007].eqiad.wmnet
* 09:36 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 09:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 09:16 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 09:05 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:04 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 09:00 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:59 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:57 dpogorzelski@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 08:55 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 08:50 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 08:49 atsuko@cumin1003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 08:49 atsukoito: depooling cirrussearch in codfw because of regression after upgrade [[phab:T431091|T431091]]
* 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts mirror1001.wikimedia.org
* 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 08:31 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:29 jmm@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: mirror1001.wikimedia.org decommissioned, removing all IPs except the asset tag one - jmm@cumin2003"
* 08:18 jmm@cumin2003: START - Cookbook sre.dns.netbox
* 08:11 jmm@cumin2003: START - Cookbook sre.hosts.decommission for hosts mirror1001.wikimedia.org
* 06:15 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 07m 18s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
== 2026-07-02 ==
* 22:55 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint1003.wikimedia.org with OS trixie
* 22:29 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint1003.wikimedia.org with reason: host reimage
* 22:23 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint1003.wikimedia.org with reason: host reimage
* 22:05 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint1003.wikimedia.org with OS trixie
* 22:03 mutante: contint1003 (zuul.wikimedia.org) - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]]
* 22:03 dzahn@cumin2002: DONE (FAIL) - Cookbook sre.hosts.downtime (exit_code=99) for 2:00:00 on zuul.wikimedia.org with reason: reimage
* 21:39 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 18s)
* 21:39 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 21:20 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards
* 21:19 sbassett: Deployed security fix for [[phab:T428829|T428829]]
* 20:58 cmooney@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003
* 20:55 cmooney@cumin1003: START - Cookbook sre.deploy.python-code homer to cumin[2002-2003].codfw.wmnet,cumin1003.eqiad.wmnet with reason: Release v0.11.2 update for new Aerleon - cmooney@cumin1003
* 20:40 arlolra@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] (duration: 12m 35s)
* 20:36 arlolra@deploy1003: cscott, arlolra: Continuing with deployment
* 20:35 bking@deploy1003: Finished deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]] (duration: 00m 20s)
* 20:35 bking@deploy1003: Started deploy [wdqs/wdqs@e8fb00c] (wcqs): [[phab:T430879|T430879]]
* 20:33 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host contint2003.wikimedia.org with OS trixie
* 20:31 arlolra@deploy1003: cscott, arlolra: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Cha
* 20:28 arlolra@deploy1003: Started scap sync-world: Backport for [[gerrit:1307218{{!}}Revert "Temporarily disable experimental ExtTagPFragment type" (T430344 T429624)]], [[gerrit:1307227{{!}}Preview: Ensure ParserMigration's handler is called to setUseParsoid (T429408)]], [[gerrit:1307223{{!}}Ensure ParserMigration is consulted if Parsoid should be used (T429408)]]
* 20:17 sbassett@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] (duration: 08m 13s)
* 20:14 dzahn@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on contint2003.wikimedia.org with reason: host reimage
* 20:13 sbassett@deploy1003: sbassett: Continuing with deployment
* 20:11 sbassett@deploy1003: sbassett: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 20:09 sbassett@deploy1003: Started scap sync-world: Backport for [[gerrit:1307196{{!}}mediawiki.action.edit.preview: Fix compat with `<button>`-buttons (T430956)]]
* 20:08 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 20:08 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 20:08 dzahn@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on contint2003.wikimedia.org with reason: host reimage
* 20:05 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs1002.eqiad.wmnet -> wcqs1003.eqiad.wmnet, repooling source-only afterwards
* 19:49 dzahn@cumin2002: START - Cookbook sre.hosts.reimage for host contint2003.wikimedia.org with OS trixie
* 19:48 mutante: contint2003 - reimaging because of [[phab:T430510|T430510]]#12067628 [[phab:T418521|T418521]]
* 18:39 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 18:17 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 18:13 bking@cumin2003: END (PASS) - Cookbook sre.wdqs.data-transfer (exit_code=0) ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards
* 17:58 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs1003.eqiad.wmnet with OS bookworm
* 17:52 jasmine@cumin2002: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-ctrl1005.eqiad.wmnet
* 17:52 jasmine@cumin2002: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-ctrl1005.eqiad.wmnet
* 17:51 jasmine@cumin2002: conftool action : set/pooled=yes:weight=10; selector: name=wikikube-ctrl1005.eqiad.wmnet
* 17:48 jasmine_: homer "cr*eqiad*" commit "Added new stacked control plane wikikube-ctrl1005"
* 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: apply
* 17:44 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: apply
* 17:31 ladsgroup@deploy1003: Finished scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] (duration: 09m 33s)
* 17:26 ladsgroup@deploy1003: ladsgroup: Continuing with deployment
* 17:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage
* 17:23 ladsgroup@deploy1003: ladsgroup: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 17:21 ladsgroup@deploy1003: Started scap sync-world: Backport for [[gerrit:1292300{{!}}etcd: Ignore test-s4 from dbctl (T427059)]]
* 17:18 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs1003.eqiad.wmnet with reason: host reimage
* 17:16 rscout@deploy1003: helmfile [eqiad] DONE helmfile.d/services/miscweb: apply
* 17:16 rscout@deploy1003: helmfile [eqiad] START helmfile.d/services/miscweb: apply
* 17:16 rscout@deploy1003: helmfile [codfw] DONE helmfile.d/services/miscweb: apply
* 17:15 rscout@deploy1003: helmfile [codfw] START helmfile.d/services/miscweb: apply
* 17:12 bking@cumin2003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 5 days, 0:00:00 on wcqs[2002-2003].codfw.wmnet,wcqs1002.eqiad.wmnet with reason: reimaging hosts
* 17:08 bd808@deploy1003: helmfile [eqiad] DONE helmfile.d/services/developer-portal: apply
* 17:08 bd808@deploy1003: helmfile [eqiad] START helmfile.d/services/developer-portal: apply
* 17:08 bd808@deploy1003: helmfile [codfw] DONE helmfile.d/services/developer-portal: apply
* 17:07 bd808@deploy1003: helmfile [codfw] START helmfile.d/services/developer-portal: apply
* 17:05 bd808@deploy1003: helmfile [staging] DONE helmfile.d/services/developer-portal: apply
* 17:05 bd808@deploy1003: helmfile [staging] START helmfile.d/services/developer-portal: apply
* 17:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 17:03 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003"
* 17:03 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "running to make sure all updates are synced - cmooney@cumin1003"
* 17:00 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs1003
* 17:00 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs1003
* 17:00 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 17:00 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs1003.eqiad.wmnet with OS bookworm
* 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Re-running - btullis@cumin1003"
* 16:58 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Re-running - btullis@cumin1003"
* 16:58 bking@cumin2003: START - Cookbook sre.wdqs.data-transfer ([[phab:T430879|T430879]], restore data on newly-reimaged host) xfer commons from wcqs2002.codfw.wmnet -> wcqs2003.codfw.wmnet, repooling source-only afterwards
* 16:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1004.eqiad.wmnet with OS bookworm
* 16:58 btullis@cumin1003: END (FAIL) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=99) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
* 16:57 tappof: bump space for prometheus k8s-aux in eqiad
* 16:55 cmooney@dns3003: END - running authdns-update
* 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 16:55 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003"
* 16:55 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to eqsin - cmooney@cumin1003"
* 16:53 cmooney@dns3003: START - running authdns-update
* 16:52 ryankemper: [ml-serve-eqiad] Cleared out 1302 failed (Evicted) pods: `kubectl -n llm delete pods --field-selector=status.phase=Failed`, freeing calico-kube-controllers from OOM crashloop (evictions were caused by disk pressure)
* 16:49 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - btullis@cumin1003"
* 16:46 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 16:39 rzl@dns1004: END - running authdns-update
* 16:37 rzl@dns1004: START - running authdns-update
* 16:36 rzl@dns1004: START - running authdns-update
* 16:35 rzl@deploy1003: Finished scap sync-world: [[phab:T416623|T416623]] (duration: 10m 19s)
* 16:34 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 16:33 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage
* 16:30 rzl@deploy1003: rzl: Continuing with deployment
* 16:28 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1004.eqiad.wmnet with reason: host reimage
* 16:26 rzl@deploy1003: rzl: [[phab:T416623|T416623]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 16:25 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 16:25 rzl@deploy1003: Started scap sync-world: [[phab:T416623|T416623]]
* 16:25 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 16:24 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/growthboo-next: sync
* 16:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/growthbook-next: sync
* 16:16 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1004.eqiad.wmnet with OS bookworm
* 16:13 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host an-test-master1003.eqiad.wmnet with OS bookworm
* 16:11 cwilliams@cumin1003: END (FAIL) - Cookbook sre.mysql.multiinstance_reboot (exit_code=99) for db-test[2001-2002].codfw.wmnet
* 16:11 cwilliams@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for db-test[2001-2002].codfw.wmnet
* 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1023: Security updates
* 16:08 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 16:08 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 16:08 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1023: Security updates
* 15:58 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage
* 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:54 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 15:54 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 15:54 btullis@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on an-test-master1003.eqiad.wmnet with reason: host reimage
* 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1023: Security updates
* 15:45 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:45 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:45 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1023: Security updates
* 15:42 btullis@cumin1003: START - Cookbook sre.hosts.reimage for host an-test-master1003.eqiad.wmnet with OS bookworm
* 15:24 moritzm: installing busybox updates from bookworm point release
* 15:20 moritzm: installing busybox updates from trixie point release
* 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool pc1021: Security updates
* 15:15 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 15:15 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 15:15 root@cumin1003: START - Cookbook sre.mysql.pool pool pc1021: Security updates
* 15:13 moritzm: installing giflib security updates
* 15:08 moritzm: installing Tomcat security updates
* 14:57 atsuko@deploy1003: helmfile [dse-k8s-codfw] DONE helmfile.d/admin 'apply'.
* 14:56 atsuko@deploy1003: helmfile [dse-k8s-codfw] START helmfile.d/admin 'apply'.
* 14:54 atsuko@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/admin 'apply'.
* 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003"
* 14:53 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003
* 14:53 atsuko@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/admin 'apply'.
* 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool pc1021: Security updates
* 14:53 root@cumin1003: END (PASS) - Cookbook sre.mysql.parsercache (exit_code=0)
* 14:53 root@cumin1003: START - Cookbook sre.mysql.parsercache
* 14:53 root@cumin1003: START - Cookbook sre.mysql.depool depool pc1021: Security updates
* 14:53 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Unblock taavi - oblivian@cumin1003
* 14:52 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Unblock taavi - oblivian@cumin1003"
* 14:46 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94711 and previous config saved to /var/cache/conftool/dbconfig/20260702-144644-fceratto.json
* 14:36 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94709 and previous config saved to /var/cache/conftool/dbconfig/20260702-143636-fceratto.json
* 14:32 moritzm: installing libdbi-perl security updates
* 14:26 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205', diff saved to https://phabricator.wikimedia.org/P94708 and previous config saved to /var/cache/conftool/dbconfig/20260702-142628-fceratto.json
* 14:16 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94707 and previous config saved to /var/cache/conftool/dbconfig/20260702-141621-fceratto.json
* 14:12 moritzm: installing rsync security updates
* 14:11 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart (exit_code=0) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 14:10 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2205 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94706 and previous config saved to /var/cache/conftool/dbconfig/20260702-140959-fceratto.json
* 14:09 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2205.codfw.wmnet with reason: Maintenance
* 14:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover
* 14:07 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:06 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1004.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 14:06 Tran: Deployed patch for [[phab:T427287|T427287]]
* 14:04 elukey@cumin1003: END (PASS) - Cookbook sre.hosts.provision (exit_code=0) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:59 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover
* 13:59 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2205: Repooling after switchover
* 13:59 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:55 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2205: Repooling after switchover
* 13:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2205 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94704 and previous config saved to /var/cache/conftool/dbconfig/20260702-135505-fceratto.json
* 13:54 moritzm: installing sed security updates
* 13:53 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:52 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2209 to s3 primary [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94703 and previous config saved to /var/cache/conftool/dbconfig/20260702-135235-fceratto.json
* 13:52 federico3: Starting s3 codfw failover from db2205 to db2209 - [[phab:T430912|T430912]]
* 13:51 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:51 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 13:48 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:47 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2209 with weight 0 [[phab:T430912|T430912]]', diff saved to https://phabricator.wikimedia.org/P94702 and previous config saved to /var/cache/conftool/dbconfig/20260702-134719-fceratto.json
* 13:47 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 24 hosts with reason: Primary switchover s3 [[phab:T430912|T430912]]
* 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:44 blake@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-pretrain: apply
* 13:44 blake@deploy1003: helmfile [codfw] START helmfile.d/services/mw-pretrain: apply
* 13:40 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 13:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 13:37 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search,name=codfw
* 13:36 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 13:36 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:34 bking@cumin2003: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 13:30 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 elukey@cumin1003: END (ERROR) - Cookbook sre.hosts.provision (exit_code=97) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:29 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:27 elukey@cumin1003: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:26 elukey@cumin1003: START - Cookbook sre.hosts.provision for host an-test-master1003.mgmt.eqiad.wmnet with chassis set policy FORCE_RESTART
* 13:25 sukhe@cumin1003: END (PASS) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=0) rolling restart_daemons on A:wikidough
* 13:23 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-psi,name=codfw
* 13:22 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search-omega,name=codfw
* 13:17 bking@cumin2003: conftool action : set/pooled=true; selector: dnsdisc=search,name=codfw
* 13:17 sukhe@puppetserver1001: conftool action : set/pooled=yes; selector: name=dns1004.wikimedia.org
* 13:12 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart (exit_code=97) rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart rolling restart_daemons on A:dnsbox and (A:dnsbox)
* 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough
* 13:11 sukhe@cumin1003: END (ERROR) - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns (exit_code=97) rolling restart_daemons on A:wikidough
* 13:11 sukhe@cumin1003: START - Cookbook sre.dns.roll-restart-reboot-wikimedia-dns rolling restart_daemons on A:wikidough
* 13:09 aude@deploy1003: Finished scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] (duration: 07m 20s)
* 13:05 aude@deploy1003: jdrewniak, aude: Continuing with deployment
* 13:04 aude@deploy1003: jdrewniak, aude: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:02 aude@deploy1003: Started scap sync-world: Backport for [[gerrit:1305773{{!}}Phase 3 Legal contact link deployments. (T430227)]]
* 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts wdqs-categories1001.eqiad.wmnet
* 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 12:19 btullis@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 12:10 jmm@dns1004: END - running authdns-update
* 12:07 jmm@dns1004: START - running authdns-update
* 11:51 btullis@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: wdqs-categories1001.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - btullis@cumin1003"
* 11:44 btullis@cumin1003: START - Cookbook sre.dns.netbox
* 11:42 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:42 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 11:41 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver1003.eqiad.wmnet
* 11:39 btullis@cumin1003: START - Cookbook sre.hosts.decommission for hosts wdqs-categories1001.eqiad.wmnet
* 11:37 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver1003.eqiad.wmnet
* 11:36 jmm@cumin2003: END (PASS) - Cookbook sre.hosts.reboot-single (exit_code=0) for host puppetserver2004.codfw.wmnet
* 11:30 jmm@cumin2003: START - Cookbook sre.hosts.reboot-single for host puppetserver2004.codfw.wmnet
* 11:29 jmm@cumin2003: END (PASS) - Cookbook sre.puppet.disable-merges (exit_code=0)
* 11:29 jmm@cumin2003: START - Cookbook sre.puppet.disable-merges
* 10:57 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2214: Repooling
* 10:49 jmm@dns1004: END - running authdns-update
* 10:47 jmm@dns1004: START - running authdns-update
* 10:31 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94698 and previous config saved to /var/cache/conftool/dbconfig/20260702-103146-fceratto.json
* 10:21 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94696 and previous config saved to /var/cache/conftool/dbconfig/20260702-102137-fceratto.json
* 10:20 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:19 fnegri@cumin1003: END (PASS) - Cookbook sre.mysql.multiinstance_reboot (exit_code=0) for clouddb1017.eqiad.wmnet
* 10:18 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.decommission (exit_code=0)
* 10:18 fceratto@cumin1003: Removing es1033 from zarcillo [[phab:T408772|T408772]]
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.decommission (exit_code=0) for hosts es1033.eqiad.wmnet
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 10:14 fceratto@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003"
* 10:14 fceratto@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: es1033.eqiad.wmnet decommissioned, removing all IPs except the asset tag one - fceratto@cumin1003"
* 10:13 fnegri@cumin1003: START - Cookbook sre.mysql.multiinstance_reboot for clouddb1017.eqiad.wmnet
* 10:12 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2214.codfw.wmnet
* 10:12 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2214.codfw.wmnet
* 10:12 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling
* 10:11 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213', diff saved to https://phabricator.wikimedia.org/P94693 and previous config saved to /var/cache/conftool/dbconfig/20260702-101130-fceratto.json
* 10:10 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:10 fceratto@cumin1003: START - Cookbook sre.dns.netbox
* 10:04 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 10:03 fceratto@cumin1003: START - Cookbook sre.hosts.decommission for hosts es1033.eqiad.wmnet
* 10:03 fceratto@cumin1003: START - Cookbook sre.mysql.decommission
* 10:01 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94691 and previous config saved to /var/cache/conftool/dbconfig/20260702-100122-fceratto.json
* 09:55 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2213 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94690 and previous config saved to /var/cache/conftool/dbconfig/20260702-095529-fceratto.json
* 09:55 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2213.codfw.wmnet with reason: Maintenance
* 09:54 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 09:53 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2213: Repooling after switchover
* 09:51 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 09:44 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2213: Repooling after switchover
* 09:39 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2213: Repooling after switchover
* 09:39 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2213 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94688 and previous config saved to /var/cache/conftool/dbconfig/20260702-093859-fceratto.json
* 09:36 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2192 to s5 primary [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94687 and previous config saved to /var/cache/conftool/dbconfig/20260702-093650-fceratto.json
* 09:36 federico3: Starting s5 codfw failover from db2213 to db2192 - [[phab:T430923|T430923]]
* 09:30 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94686 and previous config saved to /var/cache/conftool/dbconfig/20260702-093004-fceratto.json
* 09:24 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2192 with weight 0 [[phab:T430923|T430923]]', diff saved to https://phabricator.wikimedia.org/P94685 and previous config saved to /var/cache/conftool/dbconfig/20260702-092455-fceratto.json
* 09:24 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 23 hosts with reason: Primary switchover s5 [[phab:T430923|T430923]]
* 09:19 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94684 and previous config saved to /var/cache/conftool/dbconfig/20260702-091957-fceratto.json
* 09:16 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] (duration: 06m 57s)
* 09:13 moritzm: installing libgcrypt20 security updates
* 09:12 kharlan@deploy1003: kharlan: Continuing with deployment
* 09:11 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 09:09 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220', diff saved to https://phabricator.wikimedia.org/P94683 and previous config saved to /var/cache/conftool/dbconfig/20260702-090950-fceratto.json
* 09:09 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307076{{!}}SourceEditorOverlay: Re-enable buttons after non-captcha save failure (T430518)]]
* 09:03 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 09:01 kharlan@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] (duration: 07m 07s)
* 08:59 fceratto@cumin1003: dbctl commit (dc=all): 'Repooling after maintenance db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94682 and previous config saved to /var/cache/conftool/dbconfig/20260702-085942-fceratto.json
* 08:57 kharlan@deploy1003: kharlan: Continuing with deployment
* 08:56 kharlan@deploy1003: kharlan: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:54 kharlan@deploy1003: Started scap sync-world: Backport for [[gerrit:1307075{{!}}build: Update required Node version from 24.14.1 to 24.18.0]]
* 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:52 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:52 fceratto@cumin1003: dbctl commit (dc=all): 'Depooling db2220 ([[phab:T426633|T426633]])', diff saved to https://phabricator.wikimedia.org/P94681 and previous config saved to /var/cache/conftool/dbconfig/20260702-085237-fceratto.json
* 08:52 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2220.codfw.wmnet with reason: Maintenance
* 08:43 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 08:40 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group2 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 08:25 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] (duration: 11m 44s)
* 08:21 cscott@deploy1003: cscott: Continuing with deployment
* 08:16 cscott@deploy1003: cscott: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 08:14 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307059{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T387374 T430186 T430367 T430501)]], [[gerrit:1307061{{!}}Bump wikimedia/parsoid to 0.24.0-a14 (T430501)]]
* 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 08:08 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db1244: Migration of db1244.eqiad.wmnet completed
* 08:02 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:02 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-reverted' for release 'main' .
* 08:01 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] (duration: 18m 58s)
* 08:01 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:01 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-goodfaith' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-editquality-damaging' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 08:00 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-drafttopic' for release 'main' .
* 07:59 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s_services/services/datahub-next: apply
* 07:59 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:59 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:59 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2006.wikimedia.org
* 07:58 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:57 cscott@deploy1003: cscott: Continuing with deployment
* 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:56 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:56 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-draftquality' for release 'main' .
* 07:55 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:55 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articletopic' for release 'main' .
* 07:54 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2006.wikimedia.org
* 07:54 ozge@deploy1003: helmfile [ml-staging-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:54 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 07:49 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s_services/services/datahub-next: apply
* 07:44 cscott@deploy1003: cscott: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:44 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader2005.wikimedia.org
* 07:44 moritzm: installing node-lodash security updates
* 07:42 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1307058{{!}}[REST] Don't language-convert non-parsoid output; don't lookup bogus titles (T430778)]], [[gerrit:1306996{{!}}[parser] When expanding an extension tag with a title, use a new frame (T430344 T429624)]]
* 07:39 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader2005.wikimedia.org
* 07:30 cscott@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] (duration: 07m 28s)
* 07:26 cscott@deploy1003: ssastry, cscott: Continuing with deployment
* 07:25 cscott@deploy1003: ssastry, cscott: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:23 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db1244: Migration of db1244.eqiad.wmnet completed
* 07:22 cscott@deploy1003: Started scap sync-world: Backport for [[gerrit:1306985{{!}}Parsoid read views: Bump enwiki NS_MAIN desktop traffic to 100% (T430194)]]
* 07:16 wmde-fisch@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] (duration: 06m 55s)
* 07:13 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1244.eqiad.wmnet with OS trixie
* 07:11 wmde-fisch@deploy1003: wmde-fisch: Continuing with deployment
* 07:11 wmde-fisch@deploy1003: wmde-fisch: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:09 wmde-fisch@deploy1003: Started scap sync-world: Backport for [[gerrit:1306970{{!}}Fix how to check the treatment group (T415904)]], [[gerrit:1306971{{!}}Fix how to check the treatment group (T415904)]]
* 06:54 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage
* 06:50 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1244.eqiad.wmnet with reason: host reimage
* 06:38 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db1250.eqiad.wmnet with OS trixie
* 06:34 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db1244.eqiad.wmnet with OS trixie
* 06:25 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db1244: Upgrading db1244.eqiad.wmnet
* 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db1244: Upgrading db1244.eqiad.wmnet
* 06:25 cwilliams@cumin1003: dbmaint on s4@eqiad [[phab:T429893|T429893]]
* 06:25 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 06:15 marostegui@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage
* 06:14 cwilliams@dns1006: END - running authdns-update
* 06:12 cwilliams@dns1006: START - running authdns-update
* 06:11 cwilliams@dns1006: END - running authdns-update
* 06:11 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db1244 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94676 and previous config saved to /var/cache/conftool/dbconfig/20260702-061059-cwilliams.json
* 06:09 marostegui@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db1250.eqiad.wmnet with reason: host reimage
* 06:09 cwilliams@dns1006: START - running authdns-update
* 06:08 aokoth@cumin1003: END (PASS) - Cookbook sre.vrts.upgrade (exit_code=0) on VRTS host vrts1003.eqiad.wmnet
* 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db1160 to s4 primary and set section read-write [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94675 and previous config saved to /var/cache/conftool/dbconfig/20260702-060746-cwilliams.json
* 06:07 cwilliams@cumin1003: dbctl commit (dc=all): 'Set s4 eqiad as read-only for maintenance - [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94674 and previous config saved to /var/cache/conftool/dbconfig/20260702-060704-cwilliams.json
* 06:06 cezmunsta: Starting s4 eqiad failover from db1244 to db1160 - [[phab:T430817|T430817]]
* 06:04 aokoth@cumin1003: START - Cookbook sre.vrts.upgrade on VRTS host vrts1003.eqiad.wmnet
* 05:59 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db1160 with weight 0 [[phab:T430817|T430817]]', diff saved to https://phabricator.wikimedia.org/P94673 and previous config saved to /var/cache/conftool/dbconfig/20260702-055927-cwilliams.json
* 05:59 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430817|T430817]]
* 05:55 marostegui@cumin1003: START - Cookbook sre.hosts.reimage for host db1250.eqiad.wmnet with OS trixie
* 05:44 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3 days, 0:00:00 on db1250.eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]]
* 05:39 marostegui: Failover m3 (phabricator) from db1250 to db1228 - [[phab:T430158|T430158]]
* 05:32 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on db[2160,2234].codfw.wmnet,db[1217,1228,1250].eqiad.wmnet with reason: m3 master switchover [[phab:T430158|T430158]]
* 04:45 tstarling@deploy1003: Finished scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] (duration: 09m 08s)
* 04:41 tstarling@deploy1003: tstarling, reedy: Continuing with deployment
* 04:38 tstarling@deploy1003: tstarling, reedy: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 04:36 tstarling@deploy1003: Started scap sync-world: Backport for [[gerrit:1298928{{!}}CommonSettings: Set $wgScoreUseSvg = true (T49578)]]
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 59s)
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:16 ryankemper: [[phab:T429844|T429844]] [opensearch] completed `cirrussearch2111` reimage; all codfw search clusters are green, all nodes now report `OpenSearch 2.19.5`, and the temporary chi voting exclusion has been removed
* 00:57 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2111.codfw.wmnet with OS trixie
* 00:29 ryankemper: [[phab:T429844|T429844]] [opensearch] depooled codfw search-omega/search-psi discovery records to match existing codfw search depool during OpenSearch 2.19 migration
* 00:29 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage
* 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-psi,name=codfw
* 00:29 ryankemper@cumin2002: conftool action : set/pooled=false; selector: dnsdisc=search-omega,name=codfw
* 00:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2111.codfw.wmnet with reason: host reimage
* 00:01 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2111.codfw.wmnet with OS trixie
* 00:00 ryankemper: [[phab:T429844|T429844]] [opensearch] chi cluster recovered after stopping `opensearch_1@production-search-codfw` on `cirrussearch2111`
== 2026-07-01 ==
* 23:59 ryankemper: [[phab:T429844|T429844]] [opensearch] stopped `opensearch_1@production-search-codfw` on `cirrussearch2111` after chi cluster-manager election churn following `voting_config_exclusions` POST; hoping this triggers a re-election
* 23:52 cscott@deploy1003: helmfile [codfw] DONE helmfile.d/services/mw-parsoid: apply
* 23:51 cscott@deploy1003: helmfile [codfw] START helmfile.d/services/mw-parsoid: apply
* 23:51 cscott@deploy1003: helmfile [eqiad] DONE helmfile.d/services/mw-parsoid: apply
* 23:50 cscott@deploy1003: helmfile [eqiad] START helmfile.d/services/mw-parsoid: apply
* 22:37 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wcqs2003.codfw.wmnet with OS bookworm
* 22:29 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 22:13 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 22:10 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2084.codfw.wmnet with OS trixie
* 22:09 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage
* 22:03 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 22:01 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on wcqs2003.codfw.wmnet with reason: host reimage
* 21:50 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 21:43 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.move-vlan (exit_code=0) for host wcqs2003
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.network.configure-switch-interfaces (exit_code=0) for host wcqs2003
* 21:42 bking@cumin2003: START - Cookbook sre.network.configure-switch-interfaces for host wcqs2003
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.wipe-cache (exit_code=0) wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:42 bking@cumin2003: START - Cookbook sre.dns.wipe-cache wcqs2003.codfw.wmnet 45.48.192.10.in-addr.arpa 5.4.0.0.8.4.0.0.2.9.1.0.0.1.0.0.4.0.1.0.0.6.8.0.0.0.0.0.0.2.6.2.ip6.arpa on all recursors
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 21:42 bking@cumin2003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003"
* 21:42 bking@cumin2003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: Update records for host wcqs2003 - bking@cumin2003"
* 21:36 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2084.codfw.wmnet with reason: host reimage
* 21:35 bking@cumin2003: START - Cookbook sre.dns.netbox
* 21:34 bking@cumin2003: START - Cookbook sre.hosts.move-vlan for host wcqs2003
* 21:34 bking@cumin2003: START - Cookbook sre.hosts.reimage for host wcqs2003.codfw.wmnet with OS bookworm
* 21:19 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2084.codfw.wmnet with OS trixie
* 21:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2081.codfw.wmnet with OS trixie
* 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2108.codfw.wmnet with OS trixie
* 20:50 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage
* 20:45 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2081.codfw.wmnet with reason: host reimage
* 20:28 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2081.codfw.wmnet with OS trixie
* 20:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage
* 20:19 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2108.codfw.wmnet with reason: host reimage
* 19:59 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2108.codfw.wmnet with OS trixie
* 19:46 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2093.codfw.wmnet with OS trixie
* 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 19:44 jasmine@cumin2002: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002"
* 19:43 jasmine@cumin2002: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.hosts.reimage: Host reimage - jasmine@cumin2002"
* 19:42 bking@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2080.codfw.wmnet with OS trixie
* 19:28 jasmine@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 19:24 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage
* 19:18 jasmine@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on wikikube-ctrl1005.eqiad.wmnet with reason: host reimage
* 19:17 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2093.codfw.wmnet with reason: host reimage
* 19:15 bking@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage
* 19:07 bking@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2080.codfw.wmnet with reason: host reimage
* 18:57 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2093.codfw.wmnet with OS trixie
* 18:50 bking@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2080.codfw.wmnet with OS trixie
* 18:27 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 18:18 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] (duration: 09m 15s)
* 18:13 jgiannelos@deploy1003: jgiannelos, neriah: Continuing with deployment
* 18:11 jgiannelos@deploy1003: jgiannelos, neriah: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 18:09 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306950{{!}}PageBundleParserOutputConverter: Avoid revision lookup for bogus title (T430778)]], [[gerrit:1306910{{!}}PageBundleParserOutputConverter: Check for proper page before adding id/ns metadata (T430778)]]
* 17:40 jasmine@cumin2002: START - Cookbook sre.hosts.reimage for host wikikube-ctrl1005.eqiad.wmnet with OS trixie
* 16:58 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for 30 hosts
* 16:57 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 30 hosts
* 16:52 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2202.codfw.wmnet
* 16:52 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2202.codfw.wmnet
* 16:51 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt
* 16:51 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt
* 16:51 brett@cumin2002: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lvs2012.codfw.wmnet
* 16:51 brett@cumin2002: START - Cookbook sre.hosts.remove-downtime for lvs2012.codfw.wmnet
* 16:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2076.codfw.wmnet with OS trixie
* 16:49 brett: Start pybal on lvs2012 - [[phab:T429861|T429861]]
* 16:49 pt1979@cumin1003: END (ERROR) - Cookbook sre.hosts.remove-downtime (exit_code=97) for 59 hosts
* 16:48 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for 59 hosts
* 16:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2061.codfw.wmnet with OS trixie
* 16:30 dancy@deploy1003: Installation of scap version "4.271.0" completed for 2 hosts
* 16:28 dancy@deploy1003: Installing scap version "4.271.0" for 2 host(s)
* 16:23 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage
* 16:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage
* 16:18 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2076.codfw.wmnet with reason: host reimage
* 16:18 jasmine@dns1004: END - running authdns-update
* 16:16 jhancock@cumin2002: END (FAIL) - Cookbook sre.hosts.provision (exit_code=99) for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:16 jhancock@cumin2002: START - Cookbook sre.hosts.provision for host restbase2039.mgmt.codfw.wmnet with chassis set policy FORCE_RESTART
* 16:16 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2061.codfw.wmnet with reason: host reimage
* 16:15 jasmine@dns1004: START - running authdns-update
* 16:14 jasmine@dns1004: END - running authdns-update
* 16:12 jasmine@dns1004: START - running authdns-update
* 16:07 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2202.codfw.wmnet with reason: maintenance
* 16:06 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-b2-codfw,lsw1-b2-codfw IPv6,lsw1-b2-codfw.mgmt with reason: Junos upograde
* 16:00 papaul: ongoing maintenance on lsw1-b2-codfw
* 16:00 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2076.codfw.wmnet with OS trixie
* 15:59 pt1979@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt
* 15:59 pt1979@cumin1003: START - Cookbook sre.hosts.remove-downtime for lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt
* 15:57 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2061.codfw.wmnet with OS trixie
* 15:55 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) pool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:55 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node pool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:51 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 15:51 fceratto@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2220: Repooling after switchover
* 15:50 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'revscoring-articlequality' for release 'main' .
* 15:50 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 15:48 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2092.codfw.wmnet with OS trixie
* 15:41 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:40 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-test: apply
* 15:38 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 15:37 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-wikidata: apply
* 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-search: apply
* 15:35 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-search: apply
* 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 15:32 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-fr-tech: apply
* 15:30 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 15:29 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-analytics-product: apply
* 15:26 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-main: apply
* 15:25 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-main: apply
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:23 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:22 brett@cumin2002: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on lvs2012.codfw.wmnet with reason: Rack B2 maintenance - [[phab:T429861|T429861]]
* 15:21 brett: Stopping pybal on lvs2012 in preparation for codfw rack b2 maintenance - [[phab:T429861|T429861]]
* 15:20 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage
* 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:19 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/airflow-test-k8s: apply
* 15:12 _joe_: restarted manually alertmanager-irc-relay
* 15:12 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2092.codfw.wmnet with reason: host reimage
* 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] DONE helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:12 btullis@deploy1003: helmfile [dse-k8s-eqiad] START helmfile.d/dse-k8s-services/postgresql-airflow-test-k8s: apply
* 15:12 pt1979@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on lsw1-a8-codfw,lsw1-a8-codfw IPv6,lsw1-a8-codfw.mgmt with reason: Junos upograde
* 15:09 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover
* 15:07 pt1979@cumin1003: END (PASS) - Cookbook sre.k8s.pool-depool-node (exit_code=0) depool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:06 pt1979@cumin1003: START - Cookbook sre.k8s.pool-depool-node depool for host wikikube-worker[2042,2046].codfw.wmnet
* 15:02 papaul: ongoing maintenance on lsw1-a8-codfw
* 14:31 topranks: POWERING DOWN CR1-EQIAD for line card installation [[phab:T426343|T426343]]
* 14:31 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] (duration: 08m 57s)
* 14:29 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:26 dreamyjazz@deploy1003: dreamyjazz: Continuing with deployment
* 14:24 dreamyjazz@deploy1003: dreamyjazz: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 14:22 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306938{{!}}Remove group permissions definitions later in the request (T425048)]]
* 14:22 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover
* 14:16 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:15 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2220: Repooling after switchover
* 14:14 topranks: re-enable routing-engine graceful-failover on cr1-eqiad [[phab:T417873|T417873]]
* 14:13 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:13 fceratto@cumin1003: END (FAIL) - Cookbook sre.mysql.pool (exit_code=99) pool db2220: Repooling after switchover
* 14:12 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:12 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:11 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:10 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:08 dreamyjazz@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] (duration: 10m 01s)
* 14:07 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2220: Repooling after switchover
* 14:07 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2220 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94664 and previous config saved to /var/cache/conftool/dbconfig/20260701-140729-fceratto.json
* 14:06 jforrester@deploy1003: helmfile [eqiad] DONE helmfile.d/services/wikifunctions: apply
* 14:06 jforrester@deploy1003: helmfile [eqiad] START helmfile.d/services/wikifunctions: apply
* 14:06 jforrester@deploy1003: helmfile [codfw] DONE helmfile.d/services/wikifunctions: apply
* 14:05 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2159 to s7 primary [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94663 and previous config saved to /var/cache/conftool/dbconfig/20260701-140503-fceratto.json
* 14:04 jforrester@deploy1003: helmfile [codfw] START helmfile.d/services/wikifunctions: apply
* 14:04 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 14:04 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 14:04 jforrester@deploy1003: helmfile [staging] DONE helmfile.d/services/wikifunctions: apply
* 14:04 dreamyjazz@deploy1003: anzx, dreamyjazz: Continuing with deployment
* 14:04 federico3: Starting s7 codfw failover from db2220 to db2159 - [[phab:T430826|T430826]]
* 14:03 jmm@dns1004: END - running authdns-update
* 14:03 jforrester@deploy1003: helmfile [staging] START helmfile.d/services/wikifunctions: apply
* 14:03 topranks: flipping cr1-eqiad active routing-enginer back to RE0 [[phab:T417873|T417873]]
* 14:03 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on cloudsw1-c8-eqiad,cloudsw1-d5-eqiad with reason: router upgrades eqiad
* 14:01 jmm@dns1004: START - running authdns-update
* 14:00 dreamyjazz@deploy1003: anzx, dreamyjazz: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:59 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2159 with weight 0 [[phab:T430826|T430826]]', diff saved to https://phabricator.wikimedia.org/P94662 and previous config saved to /var/cache/conftool/dbconfig/20260701-135906-fceratto.json
* 13:58 dreamyjazz@deploy1003: Started scap sync-world: Backport for [[gerrit:1306456{{!}}eswikisource: add wikibooks as importsource (T430537)]], [[gerrit:1306916{{!}}Move non temporary accounts settings out TA section]], [[gerrit:1306925{{!}}Remove TA patrol rights from users on fishbowl + private (T425048)]]
* 13:57 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 28 hosts with reason: Primary switchover s7 [[phab:T430826|T430826]]
* 13:56 topranks: reboot routing-enginer RE0 on cr1-eqiad [[phab:T417873|T417873]]
* 13:48 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1006.wikimedia.org
* 13:44 atsuko@cumin2003: END (ERROR) - Cookbook sre.hosts.reimage (exit_code=97) for host cirrussearch2092.codfw.wmnet with OS trixie
* 13:43 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1006.wikimedia.org
* 13:41 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2092.codfw.wmnet with OS trixie
* 13:41 jmm@cumin2003: END (PASS) - Cookbook sre.ganeti.reboot-vm (exit_code=0) for VM urldownloader1005.wikimedia.org
* 13:37 jmm@cumin2003: START - Cookbook sre.ganeti.reboot-vm for VM urldownloader1005.wikimedia.org
* 13:37 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on pfw1-eqiad with reason: router upgrades eqiad
* 13:35 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 4:00:00 on lvs[1017-1020].eqiad.wmnet with reason: router upgrades eqiad
* 13:34 caro@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] (duration: 07m 59s)
* 13:30 caro@deploy1003: caro: Continuing with deployment
* 13:28 caro@deploy1003: caro: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:27 topranks: route-engine failover cr1-eqiad
* 13:26 caro@deploy1003: Started scap sync-world: Backport for [[gerrit:1306842{{!}}EditCheck: fix pre-save focusedAction error (T430741)]]
* 13:15 topranks: rebooting routing-engine 1 on cr1-eqiad [[phab:T417873|T417873]]
* 13:13 jgiannelos@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] (duration: 08m 29s)
* 13:13 moritzm: installing qemu security updates
* 13:11 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 13:11 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 13:09 jgiannelos@deploy1003: jgiannelos: Continuing with deployment
* 13:08 ozge@deploy1003: helmfile [ml-serve-codfw] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 13:07 jgiannelos@deploy1003: jgiannelos: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 13:06 ozge@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 13:06 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on db2214.codfw.wmnet with reason: Maintenance
* 13:05 fceratto@cumin1003: END (ERROR) - Cookbook sre.mysql.pool (exit_code=97) pool db2214: Repooling after switchover
* 13:05 jgiannelos@deploy1003: Started scap sync-world: Backport for [[gerrit:1306873{{!}}Parsoid read views: Bump enwiki traffic to 75%]]
* 13:04 fceratto@cumin1003: START - Cookbook sre.mysql.pool pool db2214: Repooling after switchover
* 13:04 fceratto@cumin1003: dbctl commit (dc=all): 'Depool db2214 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94660 and previous config saved to /var/cache/conftool/dbconfig/20260701-130413-fceratto.json
* 13:01 moritzm: installing python3.13 security updates
* 13:00 fceratto@cumin1003: dbctl commit (dc=all): 'Promote db2229 to s6 primary [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94659 and previous config saved to /var/cache/conftool/dbconfig/20260701-125959-fceratto.json
* 12:59 federico3: Starting s6 codfw failover from db2214 to db2229 - [[phab:T430814|T430814]]
* 12:57 cmooney@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 3:00:00 on 13 hosts with reason: router upgrade and line card install
* 12:51 fceratto@cumin1003: dbctl commit (dc=all): 'Set db2229 with weight 0 [[phab:T430814|T430814]]', diff saved to https://phabricator.wikimedia.org/P94658 and previous config saved to /var/cache/conftool/dbconfig/20260701-125149-fceratto.json
* 12:51 fceratto@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 21 hosts with reason: Primary switchover s6 [[phab:T430814|T430814]]
* 12:50 fceratto@cumin1003: END (PASS) - Cookbook sre.hosts.remove-downtime (exit_code=0) for db2189.codfw.wmnet
* 12:50 fceratto@cumin1003: START - Cookbook sre.hosts.remove-downtime for db2189.codfw.wmnet
* 12:42 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2100.codfw.wmnet with OS trixie
* 12:38 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2083.codfw.wmnet with OS trixie
* 12:19 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage
* 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.major-upgrade (exit_code=0)
* 12:17 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.pool (exit_code=0) pool db2240: Migration of db2240.codfw.wmnet completed
* 12:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage
* 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2083.codfw.wmnet with reason: host reimage
* 12:09 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2100.codfw.wmnet with reason: host reimage
* 12:00 topranks: drain traffic on cr1-eqiad to allow for line card install and JunOS upgrade [[phab:T426343|T426343]]
* 11:52 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2083.codfw.wmnet with OS trixie
* 11:50 cmooney@dns2005: END - running authdns-update
* 11:49 cmooney@dns2005: START - running authdns-update
* 11:48 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2100.codfw.wmnet with OS trixie
* 11:40 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/zotero: apply
* 11:40 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/zotero: apply
* 11:36 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/zotero: apply
* 11:36 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/zotero: apply
* 11:31 cwilliams@cumin1003: START - Cookbook sre.mysql.pool pool db2240: Migration of db2240.codfw.wmnet completed
* 11:30 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/zotero: apply
* 11:28 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/zotero: apply
* 11:27 mvolz@deploy1003: helmfile [eqiad] DONE helmfile.d/services/citoid: apply
* 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:27 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:27 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:27 mvolz@deploy1003: helmfile [eqiad] START helmfile.d/services/citoid: apply
* 11:23 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 11:21 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host db2240.codfw.wmnet with OS trixie
* 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.dns.netbox (exit_code=0)
* 11:20 cmooney@cumin1003: END (PASS) - Cookbook sre.puppet.sync-netbox-hiera (exit_code=0) generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:17 cmooney@cumin1003: START - Cookbook sre.puppet.sync-netbox-hiera generate netbox hiera data: "Triggered by cookbooks.sre.dns.netbox: add new link IP dns for trasnport circuits to magru - cmooney@cumin1003"
* 11:16 mvolz@deploy1003: helmfile [codfw] DONE helmfile.d/services/citoid: apply
* 11:16 mvolz@deploy1003: helmfile [codfw] START helmfile.d/services/citoid: apply
* 11:15 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2086.codfw.wmnet with OS trixie
* 11:14 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2106.codfw.wmnet with OS trixie
* 11:14 mvolz@deploy1003: helmfile [staging] DONE helmfile.d/services/citoid: apply
* 11:13 mvolz@deploy1003: helmfile [staging] START helmfile.d/services/citoid: apply
* 11:12 cmooney@cumin1003: START - Cookbook sre.dns.netbox
* 11:09 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2115.codfw.wmnet with OS trixie
* 11:04 cwilliams@cumin1003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on db2240.codfw.wmnet with reason: host reimage
* 11:00 cwilliams@cumin1003: START - Cookbook sre.hosts.downtime for 2:00:00 on db2240.codfw.wmnet with reason: host reimage
* 10:53 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage
* 10:49 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage
* 10:44 atsuko@cumin2003: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage
* 10:44 cwilliams@cumin1003: START - Cookbook sre.hosts.reimage for host db2240.codfw.wmnet with OS trixie
* 10:44 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2086.codfw.wmnet with reason: host reimage
* 10:42 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2106.codfw.wmnet with reason: host reimage
* 10:41 cwilliams@cumin1003: END (PASS) - Cookbook sre.mysql.depool (exit_code=0) depool db2240: Upgrading db2240.codfw.wmnet
* 10:41 cwilliams@cumin1003: START - Cookbook sre.mysql.depool depool db2240: Upgrading db2240.codfw.wmnet
* 10:41 cwilliams@cumin1003: dbmaint on s4@codfw [[phab:T429893|T429893]]
* 10:40 cwilliams@cumin1003: START - Cookbook sre.mysql.major-upgrade
* 10:39 atsuko@cumin2003: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2115.codfw.wmnet with reason: host reimage
* 10:27 cwilliams@cumin1003: dbctl commit (dc=all): 'Depool db2240 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94653 and previous config saved to /var/cache/conftool/dbconfig/20260701-102658-cwilliams.json
* 10:26 moritzm: installing nginx security updates
* 10:26 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2086.codfw.wmnet with OS trixie
* 10:23 cwilliams@cumin1003: dbctl commit (dc=all): 'Promote db2179 to s4 primary [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94652 and previous config saved to /var/cache/conftool/dbconfig/20260701-102356-cwilliams.json
* 10:23 cezmunsta: Starting s4 codfw failover from db2240 to db2179 - [[phab:T430127|T430127]]
* 10:23 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2106.codfw.wmnet with OS trixie
* 10:20 atsuko@cumin2003: START - Cookbook sre.hosts.reimage for host cirrussearch2115.codfw.wmnet with OS trixie
* 10:15 cwilliams@cumin1003: dbctl commit (dc=all): 'Set db2179 with weight 0 [[phab:T430127|T430127]]', diff saved to https://phabricator.wikimedia.org/P94651 and previous config saved to /var/cache/conftool/dbconfig/20260701-101531-cwilliams.json
* 10:15 cwilliams@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 40 hosts with reason: Primary switchover s4 [[phab:T430127|T430127]]
* 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003"
* 09:56 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003
* 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template (take 2) - oblivian@cumin1003
* 09:55 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template (take 2) - oblivian@cumin1003"
* 09:51 bwojtowicz@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'llm' for release 'main' .
* 09:39 mszwarc@deploy1003: Synchronized private/SuggestedInvestigationsSignals/SuggestedInvestigationsSignal4n.php: Update SI signal 4n (duration: 06m 08s)
* 09:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 09:21 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 09:14 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 09:14 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 09:02 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 09:02 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 08:54 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group0 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 08:38 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 08:38 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 08:36 aklapper@deploy1003: rebuilt and synchronized wikiversions files: group1 to 1.47.0-wmf.9 refs [[phab:T423918|T423918]]
* 08:21 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1002.wikimedia.org
* 08:21 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] (duration: 36m 11s)
* 08:15 filippo@cumin1003: conftool action : set/pooled=yes; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 08:09 mszwarc@deploy1003: mszwarc, abi: Continuing with deployment
* 08:03 mszwarc@deploy1003: mszwarc, abi: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:55 filippo@cumin1003: conftool action : set/pooled=no; selector: service=dumps-nfs,name=clouddumps1001.wikimedia.org
* 07:51 gkyziridis@deploy1003: helmfile [ml-serve-eqiad] Ran 'sync' command on namespace 'experimental' for release 'main' .
* 07:45 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306850{{!}}ULS rewrite: change description key in EmptySearchEntrypoint (T429882)]]
* 07:30 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050] (duration: 00m 22s)
* 07:29 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train 2nd try [analytics/refinery@410f2050]
* 07:28 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050] (duration: 01m 59s)
* 07:28 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] (duration: 07m 19s)
* 07:26 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (thin): Regular analytics weekly train THIN [analytics/refinery@410f2050]
* 07:26 aqu@deploy1003: Finished deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050] (duration: 04m 32s)
* 07:24 mszwarc@deploy1003: wmde-fisch, mszwarc: Continuing with deployment
* 07:23 mszwarc@deploy1003: wmde-fisch, mszwarc: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:21 aqu@deploy1003: Started deploy [analytics/refinery@410f205]: Regular analytics weekly train [analytics/refinery@410f2050]
* 07:21 aqu@deploy1003: Finished deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050] (duration: 02m 01s)
* 07:20 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306710{{!}}Fix async loading in footnote click interaction experiment (T415904)]], [[gerrit:1306711{{!}}Fix async loading in footnote click interaction experiment (T415904)]]
* 07:19 aqu@deploy1003: Started deploy [analytics/refinery@410f205] (hadoop-test): Regular analytics weekly train TEST [analytics/refinery@410f2050]
* 07:13 mszwarc@deploy1003: Finished scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] (duration: 09m 13s)
* 07:09 mszwarc@deploy1003: mszwarc, chlod, revi: Continuing with deployment
* 07:06 mszwarc@deploy1003: mszwarc, chlod, revi: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]] synced to the testservers (see https://wikitech.wikimedia.org/wiki/Mwdebug). Changes can now be verified there.
* 07:04 mszwarc@deploy1003: Started scap sync-world: Backport for [[gerrit:1306304{{!}}frwiki: change to Wikipedia 25 logo (T430409)]], [[gerrit:1306221{{!}}Temporarily change plwiki tagline for 1.7M articles (T430512)]], [[gerrit:1306649{{!}}CommonSettings: add Ombuds to wgWMCGlobalGroupToRateLimitClass (T430641)]]
* 06:55 elukey: upgrade all trixie hosts to pywmflib 3.0 - [[phab:T430552|T430552]]
* 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:43 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:43 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:42 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Fix template - oblivian@cumin1003
* 06:41 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Fix template - oblivian@cumin1003"
* 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.hiddenparma (exit_code=0) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:35 oblivian@cumin1003: END (PASS) - Cookbook sre.deploy.python-code (exit_code=0) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:34 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:31 jmm@cumin2003: DONE (PASS) - Cookbook sre.idm.logout (exit_code=0) Logging Niharika29 out of all services on: 2453 hosts
* 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.hiddenparma (exit_code=99) Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:30 oblivian@cumin1003: END (FAIL) - Cookbook sre.deploy.python-code (exit_code=99) hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.python-code hiddenparma to alert[1002,2002].wikimedia.org with reason: Various improvements - oblivian@cumin1003
* 06:30 oblivian@cumin1003: START - Cookbook sre.deploy.hiddenparma Hiddenparma deployment to the alerting hosts with reason: "Various improvements - oblivian@cumin1003"
* 06:01 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2109.codfw.wmnet with OS trixie
* 05:45 marostegui@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2 days, 0:00:00 on es1039.eqiad.wmnet with reason: issues
* 05:41 marostegui@cumin1003: conftool action : set/weight=100; selector: name=clouddb1027.eqiad.wmnet
* 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2068.codfw.wmnet with OS trixie
* 05:40 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage
* 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s2
* 05:40 marostegui@cumin1003: conftool action : set/pooled=yes; selector: name=clouddb1027.eqiad.wmnet,service=s7
* 05:36 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2109.codfw.wmnet with reason: host reimage
* 05:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage
* 05:16 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2109.codfw.wmnet with OS trixie
* 05:15 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2068.codfw.wmnet with reason: host reimage
* 05:09 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2067.codfw.wmnet with OS trixie
* 04:56 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2068.codfw.wmnet with OS trixie
* 04:49 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage
* 04:45 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2067.codfw.wmnet with reason: host reimage
* 04:27 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2067.codfw.wmnet with OS trixie
* 03:47 slyngshede@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1 day, 0:00:00 on es1039.eqiad.wmnet with reason: Hardware crash
* 03:21 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2107.codfw.wmnet with OS trixie
* 02:59 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage
* 02:55 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2085.codfw.wmnet with OS trixie
* 02:51 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2072.codfw.wmnet with OS trixie
* 02:51 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2107.codfw.wmnet with reason: host reimage
* 02:35 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage
* 02:31 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2107.codfw.wmnet with OS trixie
* 02:30 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage
* 02:26 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2085.codfw.wmnet with reason: host reimage
* 02:22 ryankemper@cumin2002: START - Cookbook sre.hosts.downtime for 2:00:00 on cirrussearch2072.codfw.wmnet with reason: host reimage
* 02:09 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2085.codfw.wmnet with OS trixie
* 02:07 mwpresync@deploy1003: Finished scap build-images: Publishing wmf/next image (duration: 06m 54s)
* 02:03 ryankemper@cumin2002: START - Cookbook sre.hosts.reimage for host cirrussearch2072.codfw.wmnet with OS trixie
* 02:00 mwpresync@deploy1003: Started scap build-images: Publishing wmf/next image
* 01:07 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad back to read-write - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94649 and previous config saved to /var/cache/conftool/dbconfig/20260701-010716-ladsgroup.json
* 01:05 ladsgroup@dns1004: END - running authdns-update
* 01:05 ladsgroup@cumin1003: dbctl commit (dc=all): 'Depool es1039 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94648 and previous config saved to /var/cache/conftool/dbconfig/20260701-010551-ladsgroup.json
* 01:03 ladsgroup@dns1004: START - running authdns-update
* 01:00 ladsgroup@cumin1003: dbctl commit (dc=all): 'Promote es1035 to es7 primary [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94647 and previous config saved to /var/cache/conftool/dbconfig/20260701-010002-ladsgroup.json
* 00:58 Amir1: Starting es7 eqiad failover from es1039 to es1035 - [[phab:T430765|T430765]]
* 00:53 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es1035 with weight 0 [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94646 and previous config saved to /var/cache/conftool/dbconfig/20260701-005329-ladsgroup.json
* 00:53 ladsgroup@cumin1003: DONE (PASS) - Cookbook sre.hosts.downtime (exit_code=0) for 1:00:00 on 9 hosts with reason: Primary switchover es7 [[phab:T430765|T430765]]
* 00:42 ladsgroup@cumin1003: dbctl commit (dc=all): 'Set es7 eqiad as read-only for maintenance - [[phab:T430765|T430765]]', diff saved to https://phabricator.wikimedia.org/P94645 and previous config saved to /var/cache/conftool/dbconfig/20260701-004221-ladsgroup.json
* 00:20 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2102.codfw.wmnet with OS trixie
* 00:15 ryankemper@cumin2002: END (PASS) - Cookbook sre.hosts.reimage (exit_code=0) for host cirrussearch2103.codfw.wmnet with OS trixie
* 00:05 dr0ptp4kt: DEPLOYED Refinery at {{Gerrit|4e7a2b32}} for changes: pageview allowlist {{Gerrit|1305158}} (+min.wikiquote) {{Gerrit|1305162}} (+bol.wikipedia), {{Gerrit|1305156}} (+isv.wikipedia); {{Gerrit|1305980}} (pv allowlist -api.wikimedia, sqoop +isvwiki); sqoop {{Gerrit|1295064}} (+globalimagelinks) {{Gerrit|1295069}} (+filerevision) using scap, then deployed onto HDFS (manual copyToLocal required additionally)
== Other archives ==
See [[Server Admin Log/Archives]].
<noinclude>
[[Category:SAL]]
[[Category:Operations]]
</noinclude>
pokig9z7worxtljlr3vr3eorgd42qyf
Etcd/Main cluster
0
441014
2445248
2445035
2026-08-09T03:45:35Z
Quiddity
1884
lang="bash"
2445248
wikitext
text/x-wiki
The '''main etcd cluster''' is the [[Etcd|Etcd cluster]] used as a state management system for the WMF production cluster. It is operated by SRE Service Ops under the [[SLO/etcd_main_cluster|etcd main cluster SLO]].
== Usage in production ==
More and more systems depend on etcd for retrieving state information. All current uses are listed in the table below
{| class="wikitable"
!software
!use
!connection
!interval
!failure mode
|-
|pybal/LVS
|retrieve LB pools servers lists, weights, state
|custom python/twisted, host only
|watch
|will keep working until restart
|-
|varnish/traffic
|retrieve list of backend servers; retrieve VCL fragments (requestctl)
|confd (watch)
|watch
|will keep working
|-
|gdnsd/auth dns
|Write admin state files for discovery.wmnet records
|confd
|watch
|will keep working
|-
|scap/deployment
|Dsh lists
|confd
|60 s
|will keep working
|-
|[[MediaWikiEtcdConfig|MediaWiki]]
|fetch some config variables
|PHP connection, request at intervals
|10 s
|PHP-FPM workloads will keep working until restart, [[Mw-cron jobs]] will fail to start
|-
|Icinga servers
|Update a local cache of the last modified index to be used by other checks
|cURL
|30 s
|the checks will use stale data for comparison
|-
|Spicerack
|distributed locking for cookbook executions
|conftool as a Python library
|n.a.
|cookbooks can be run without acquiring the lock if needed
|-
|[[MariaDB/Zarcillo|Zarcillo]]
|reads candidate master host list
|Python etcd client
|2 m
|graceful/keeps working
|-
|[[Conftool2git]]
|combines audit log events and etcd backend state to mirror history to git
|conftool as a Python library
|n.a.
|audit log events will not trigger updates to the git mirror
|-
|[[Requestctl]]
|web and command line interface for managing WAF rules
|conftool as a Python library
|n.a.
|responders are unable to manage / modify WAF rules
|}
In a failure, all systems will become unable to modify any configuration it derives from etcd, but they will keep working. Only a subset of those will survive a service restart though.
== Architecture ==
The main cluster is composed of two separated sub-clusters: the "codfw.wmnet" and "eqiad.wmnet" ones (creatively name after the datacenters they're located in) that are not connected via RAFT consensus, but via replication, so that there is always a master cluster and a replica one.
=== Consistency ===
For reads that don't require sub-second consistency cluster-wide, reading from the replica cluster is acceptable. If replication breaks, this will page opsens that will be able to correct the issue quickly enough (worst case scenario, by pointing clients to the master dc), All writes should go to the master datacenter; we ensure that the replica cluster is in read-only mode for remote clients to avoid issues.
=== Replication ===
Replication works using [https://github.com/wikimedia/operations-software-etcd-mirror etcdmirror] - a pretty raw software we wrote internally that allows replicating from one cluster to another mangling key prefixes. This is supposed to offer the functionality that <code>etcdctl mirror-maker</code> provides on etcd 3 to etcd 2 clusters.
Etcdmirror runs from one machine on the replica cluster (see the <code>profile::etcd::replication::active</code> hiera key). It reads the etcd index to replicate from in <code>/__replication/$destination_prefix</code> (or, if <code>$destination_prefix</code> is the root of the replica cluster keyspace <code>/</code>, to <code>/__replication/__ROOT__</code>), issues a recursive watch request to the source cluster starting at the recorded index, and then recursively replicates every write that happens under <code>$source_prefix</code> in the source cluster.
As of April 2024 ([[phab:T358636|T358636]]), we're replicating nearly the entire keyspace (i.e., <code>/</code> to <code>/</code>), including <code>/conftool</code> ([[Conftool|conftool]] state) and <code>/spicerack</code> ([[Spicerack|spicerack]] lock state). One notable exception is <code>/spicerack/locks/etcd</code> which contains short-lived [https://github.com/jplana/python-etcd python-etcd] lock state that isn't meaningful outside of the source cluster, and is thus ignored by replication.
The logs produced by etcdmirror are pretty verbose, detailing each replication event and any errors encountered should anything go wrong.
==== Recovering from replication failures ====
In the event that etcdmirror fails (indicated by the EtcdReplicationDown alert), it should be safe to try restarting the systemd unit if logs suggest a transient issue - e.g., connectivity to the source cluster.
However, etcdmirror is very strict when applying operations to the destination cluster and will fail as soon as any inconsistency is found (even just in the original value of a key) or if the lag is large enough that we're losing etcd events (i.e., when the latest event we've been able to replicate falls outside the 1000 event retention window at the source etcd cluster; see the note in [https://etcd.io/docs/v2.3/api/#waiting-for-a-change this section] of the etcd API docs).
In such a case, you will need to do a full reload. To do that, you need to launch etcdmirror with identical arguments to those used by the systemd unit, but adding the <code>--reload</code> switch. There is a shell script available in <code>/usr/local/sbin</code> on the replication host, which does this for you (look for <code>reload-etdmirror</code>). Once the reload is complete (look for "Starting replication at" in the logs), you can stop your manual invocation of etcdmirror and restart the systemd unit.
{{warning|Beware: doing so will <u>ERASE ALL DATA</u> on the destination cluster, so do that with extreme caution.}}
=== Individual cluster configuration ===
We decided to proxy external connections to etcd via an nginx proxy that handles TLS and HTTP authentication and should be fully compliant with etcd's own behaviour. The reason for this is that the builtin authentication gives a severe performance hit to etcd, and that our TLS configuration for nginx is much better than what etcd itself offers. It also gives us the ability to switch on/off the read-only status of a cluster by flipping a switch in puppet. I don't know of any way to do this with the standard etcd mechanism without actually removing users and/or roles, a slow process that is hard to automate/puppetize.
Instead, what happens is that on every cluster member we have an etcd instance listening for client connections on <nowiki>https://$fqdn:2379</nowiki> with no authentication, but inaccessible to external connections (firewall rules). So local clients, such as etcdmirror, can write to it unauthenticated. At the same time, etcd advertises <nowiki>https://$fqdn:4001</nowiki> as its client URL, which is also where nginx is listening for external connections and enforces authentication as well.
{{note|This can be surprising in exceptional cases when you need to directly modify keys in a cluster where nginx is enforcing read-only mode. If you attempt to do so with <code>etcdctl --endpoints <nowiki>https://$fqdn:2379</nowiki></code>, it will still use the advertised URLs and writes will be rejected by nginx.}}
To work around this, you can use <code>curl</code> to issue the equivalent [https://etcd.io/docs/v2.3/api etcd v2 API] calls against <nowiki>https://$fqdn:2379</nowiki>. Again, directly modifying keys is an exceptional operation, so consider getting your commands reviewed by a peer.
The TLS certificates used by etcd - i.e., for peer-peer and client-facing connections (with the sole client being the collocated proxy) - and proxy are managed by our cfssl-based [[PKI]] (see {{phabricator|T352245}}). The former uses a [[PKI/CA_Operations#Intermediate_Certificates|custom intermediate]] to enable TLS client auth for peer-peer connections.
== Operations ==
For the most part, you can refer to what is written in [[Etcd]], but there are a few more operations regarding replication that are not covered there.
=== Primary cluster switchover ===
This procedure was refreshed in August 2026 as part of {{phabricator|T433554}}. See {{phabricator|T166552}} for historical prior art (2017).
{{Note|type=warn|text=This is a '''rarely used and high-risk''' procedure. Carefully review each step for accuracy and relevance before proceeding.}}
Assuming all patches have been prepared ahead of time, the following procedure incurs an etcd read-only period of '''15 - 20 minutes'''. During that time:
* Mutating operations by <code>confctl</code>, <code>dbctl</code>, and <code>requestctl</code> will fail (reads will succeed); note that this includes <code>conftool-sync</code> during <code>puppet-merge</code> runs for patches that update puppet's <code>conftool-data</code>.
* Cookbooks will be unable to acquire or release Spicerack locks, which are also backed by etcd.
Given this disruption to normal maintenance operations, it is important to communicate the planned switchover to SRE in advance, ideally at least 24 hours ahead of the announced maintenance window, and communicate status proactively in IRC while the process is in flight.
The following procedure contains embedded examples from {{phabricator|T433554}}, which switched the primary cluster from eqiad to codfw. When moving in the opposite direction, swap the data centers and / or hosts accordingly in each step.
Unless otherwise noted, commands are run on a cluster-management host. The procedure is broken down into three phases, where the '''Switchover''' phase is the disruptive one (read-only). The '''Preparation''' and '''Cleanup''' phases should happen shortly before and shortly, respectively.
==== Preparation ====
1. Reduce the TTL for conftool (R/W) etcd client SRV records to 10 seconds ([[DNS#Deploying_DNS_changes]]).
* Example: {{gerrit|1319194}}
2. [[Alertmanager#Silences_&_acknowledgements|Silence]] the EtcdReplicationDown alert shortly before work begins.
==== Switchover ====
1. Begin read-only in the etcd cluster we are switching from. Disruption begins.
* Example: {{gerrit|1319191}}. Deployed with <code>sudo cumin A:conf-eqiad 'run-puppet-agent'</code>
2. Verify read-only
* Attempting to depool a suitable host should fail (e.g., <code>sudo confctl select "name=$FDQN" set/pooled=no</code>).
3. Disable puppet on the current and new replication hosts.
* Example: <code>sudo cumin 'conf2005.codfw.wmnet,conf1008.eqiad.wmnet' 'disable-puppet "etcd replication switchover"'</code>
4. Merge a puppet patch that disables etcd replication in the cluster we are switching to and enables it in the cluster we are switching from (i.e., affecting the two hosts on which you just disabled puppet).
* Example: {{gerrit|1319192}}
5. Run puppet-agent on the current replication host in the cluster we are switching to (replication stops).
* Example:
** (on conf2005.codfw.wmnet) <code>sudo run-puppet-agent -e "etcd replication switchover"</code>
** (on conf2005.codfw.wmnet) confirm that <code>etcdmirror--eqiad-wmnet.service</code> has terminated (this may take up to 60 seconds)
6. Set the replication index in the cluster we are switching from.
* Use the python script in [https://phabricator.wikimedia.org/P95724 P95724], which can be invoked from any member of the cluster we're switching from (e.g., the new replication host). Note that you will be prompted to approve the etcd write.
* Example:
** (on conf1008.eqiad.wmnet) <code>python3 set_mirror_index.py --protocol https --old-replica conf2005.codfw.wmnet:4001 --new-replica $(hostname -f):2379 --prefix __ROOT__</code>
** Note: the prefix is <code>__ROOT__</code> because we are replicating the full keyspace; see [[#Replication]] above.
7. Run puppet-agent on the new replication host in the cluster we are switching from (replication starts).
* Example:
** (on conf1008.eqiad.wmnet) <code>sudo run-puppet-agent -e "etcd replication switchover"</code>
** (on conf1008.eqiad.wmnet) confirm that <code>etcdmirror--codfw-wmnet.service</code> has started
8. Test replication via a local mutation on any member of the cluster we are switching to.
* Example:
** (on conf1008.eqiad.wmnet) Monitor replication progress: <code>journalctl -f -u etcdmirror--codfw-wmnet.service</code>
** (on conf2005.codfw.wmnet or any other codfw cluster member) Write (and then delete) a suitable test key:
*** <code>curl https://$(hostname -f):2379/v2/keys/test -XPUT -d value=""</code>
*** <code>curl https://$(hostname -f):2379/v2/keys/test -XDELETE</code>
9. Switch the conftool (R/W) etcd client SRV records to the cluster we are switching to ([[DNS#Deploying_DNS_changes]]).
* Example: {{gerrit|1319195}}
10. Begin read-write in the etcd cluster we are switching to.
* Example: {{gerrit|1319193}}. Deployed with <code>sudo cumin A:conf-codfw 'run-puppet-agent'</code>
11. Verify read-write
* Attempting to depool a suitable host should succeed, e.g., <code>sudo confctl select "name=$FQDN" set/pooled=no</code>.
* Remember to repool the host (e.g., <code>sudo confctl select "name=$FQDN" set/pooled=yes</code>).
12. Restart hiddenparma ([[Requestctl]]), since it will have cached connections toward the cluster we are switching from, which is now read-only - i.e., subsequent edits will fail.
* (on alert1002.wikimedia.org, alert2002.wikimedia.org) <code>sudo systemctl restart hiddenparma.service</code>
13. Restart additional '''read-only''' services that happen to track the conftool-specific SRV records. Since these are read-only, this is far less urgent than anything above (i.e., they will continue to function even if still connected to the cluster we have switched from).
* [[Conftool2git]] - <code>sudo systemctl restart conftool2git.service</code> on the <code>profile::conftool2git::active_host</code>.
* [[MariaDB/Zarcillo]] - A service running in the aux-k8s-eqiad cluster. Coordinate with Data Persistence (i.e., ask them to restart it).
==== Cleanup ====
1. Delete the EtcdReplicationDown silence.
* If the alert is still firing in the cluster we are switching to (i.e., where etcd-mirror was running previously), it means the ops prometheus host(s) in that DC need a puppet-agent run (you can just wait for this to happen organically before deleting the silence).
2. Restore the TTL for conftool (R/W) etcd client SRV records to 5 minutes ([[DNS#Deploying_DNS_changes]]).
* Example: {{gerrit|1319196}}
=== Depool etcd client traffic from a cluster ===
{{Note|type=warn|text=If you intend to depool the current primary cluster, you first need to perform a switchover as described above.}}
In what follows, the ''target'' cluster refers to the cluster that will be depooled. For a recent example of this procedure in use, see {{phabricator|T430909}}. As always, '''reality may have changed''' since the time of this writing, so review the procedure closely for steps that seem unclear or potentially out of date before starting, seeking clarification from your colleagues as needed. Finally, please <code>!log</code> extensively as you perform these steps.
==== Move PyBals ====
As of mid-2026, PyBal is still used in core DCs. PyBals in a given DC connect to single static etcd cluster node in that same DC (unless depooled, as we're doing here), specified via the <code>profile::pybal::config_host</code> hiera key in puppet.
* Locate the instance(s) of this hiera key in puppet that configure PyBals to use an etcd node in the target cluster. In practice, this will be one of either <code>hieradata/role/codfw/lvs/balancer.yaml</code> or <code>hieradata/role/eqiad/lvs/balancer.yaml</code> depending on target cluster.
* Switch the value of <code>profile::pybal::config_host</code> to an etcd node in the ''other'' cluster. Example: {{Gerrit|1308115}}.
* '''Coordinate with Traffic''': For each of the LVS hosts affected by your change, run puppet-agent and then restart PyBal.
** See below for details on how to apply this change to a given LVS host. Additional resources can be found in [[LVS#Configure_the_load_balancers]] (e.g., using <code>cumin</code> or the <code>sre.loadbalancer.restart-pybal</code> cookbook).
** Note that you will typically be touching ''only'' PyBals in a single (core) DC, including all 4 of the secondary, low-traffic, high-traffic1, and high-traffic2 instances therein.
===== Point a PyBal host to a different etcd host =====
{{Note|type=info|text=This assumes that the relevant changes have been merged on Gerrit and in Puppet.}}
* Run puppet on the LVS host to move:
<code>$ sudo run-puppet-agent</code>
* Restart PyBal:
<code>$ sudo systemctl restart pybal.service</code>
* Verify the service and check the journal for potential problems:
<syntaxhighlight lang="bash">
$ sudo journalctl -ru pybal.service
-- Journal begins at Wed 2026-06-03 04:17:42 UTC, ends at Wed 2026-08-05 15:55:11 UTC. --
Aug 05 15:54:08 lvs1017 pybal[3089303]: [config-etcd] INFO: connected to etcd://conf2006.codfw.wmnet:4001/conftool/v1/pools/eqiad/cache_text/cdn/ /eqiad/kubernetes-staging/kubemaster/
Aug 05 15:54:08 lvs1017 pybal[3089303]: [config-etcd] INFO: connected to etcd://conf2006.codfw.wmnet:4001/conftool/v1/pools/eqiad/ncredir/nginx/
...
</syntaxhighlight>
* Check the output of one (or more) service pool(s), to ensure that they are in the expected state(s):
<syntaxhighlight lang="bash">$ curl localhost:9090/pools
textlb6_80
ncredirlb_80
gerritlb_29418
textlb_80
$ curl localhost:9090/pools/textlb6_80
cp1100.eqiad.wmnet: enabled/up/pooled
cp1102.eqiad.wmnet: enabled/up/pooled
[...]
</syntaxhighlight>
==== Move other clients ====
Other clients (confd, MediaWiki, etc.) discover etcd cluster nodes via DNS SRV records. We use a simple geo-mapping technique that creates SRV records in the <code>wmnet</code> zone for each core and caching DC, where the constituent hosts of the SRV record point to the nearest core-DC etcd cluster. These per-DC SRV records are then used by the clients located there. For example, although ulsfo contains no etcd cluster (it's a caching DC), clients located there may use the <code>_etcd-client-ssl._tcp.ulsfo.wmnet</code> SRV record, which resolves to the etcd cluster nodes in codfw.
* Change the <code>_etcd._tcp</code> and <code>_etcd-client-ssl._tcp</code> SRV records in the <code>wmnet</code> zone corresponding to DCs mapped to the target etcd cluster. Example: {{Gerrit|1308114}}.
{{Note|type=warn|text=There is also a single SRV record in the <code>wikimedia.org</code> zone (see [https://gerrit.wikimedia.org/g/operations/dns/+/refs/heads/master/templates/wikimedia.org templates/wikimedia.org]). If you are depooling the cluster referenced there, you will need to update these records as well.}}
* Merge and deploy your change using <code>authdns-update</code>. See [[DNS#Deploying_DNS_changes]].
* You should see traffic start to shift away from the target cluster shortly after. See the [https://grafana.wikimedia.org/d/tTE9nvdMk/etcd?orgId=1 etcd grafana dashbord]. Note that clients shifting at this stage are limited to those that do not cache connections (so, basically just MediaWiki).
* Wait at least 5 minutes for the SRV TTL to pass, or explicitly clear the recursor caches for each of the SRV records updated using the <code>sre.dns.wipe-cache</code> cookbook. After this is complete, we can begin restarting clients that ''do'' cache connections.
* Rolling restart confd in all affected DCs. For example, if you are targeting the codfw cluster, and have thus updated the SRV records used by eqsin, codfw, and ulsfo clients:
<syntaxhighlight lang="bash">
sudo cumin -b 16 -s 10 -p 95 'P{C:confd} and (A:eqsin or A:codfw or A:ulsfo)' 'systemctl restart confd'
</syntaxhighlight>
{{Note|type=info|text=If you have changed the SRV record in the <code>wikimedia.org</code> zone, the clients of which are not limited to a specific subset of DCs, you can use <code>'P{C:confd} and P{F:domain = wikimedia.org}'</code> to target the affected confd instances that also require restart.}}
* Note that this uses a completion percentage below 100 in order for the restart not to abort on the first error (i.e., just in case one of the many affected hosts is transiently unreachange). If you do observe any hosts fail, investigate and whack-a-mole them individually as-needed.
* Restart <code>navtiming.service</code> on the webperf host in the affected core DC (e.g., webperf2003 if you are targeting the codfw cluster).
* '''Coordinate with Traffic''': Restart Liberica daemons in all affected DCs (note: it's really just the control-plane we want to restart, but there's not a separate cookbook for that at this time). For example, to restart Liberica in ulsfo:
<syntaxhighlight lang="bash">
sudo cookbook sre.loadbalancer.upgrade -t TXXXXXX --seamless --alias liberica-ulsfo --reason 'Clear control-plane connections to etcd' restart
</syntaxhighlight>
* Once all of these steps are complete, etcd nodes in the target cluster should no longer receive traffic from clients.
** Check <code>/var/log/nginx/etcd_access.log</code> and confirm that only <code>/metrics</code> collection remains.
** Check <code>sudo ss -apn | grep :4001</code> to confirm that no established TCP connections remain (note this will show nginx as the owning process, which terminates TLS on behalf of etcd clients).
{{Note|type=info|text=Although it will not appear in the nginx access logs, you may see a trickle of etcd requests in the depooled cluster in the grafana dashboard (mainly PUT / DELETE). This is the local etcdmirror instance replicating from the other cluster - see [[Etcd/Main_cluster#Replication]] above.}}
==== Repooling ====
Repooling the cluster follows exactly the same procedure, the only difference being that you are now '''reverting''' your earlier puppet and DNS patched (i.e., you apply those changes the same exact way).
=== Reimage cluster ===
This procedure was refreshed in August 2026 as part of {{phabricator|T428495}}.
{{Note|type=warn|content=Be aware that '''this might not reflect current reality'''! Various details might have changed ... actually everything might have changed. So please use this as a starting template and DOUBLE CHECK EVERY STEP BEFORE YOU EVEN START!}}
==== Preparation ====
* '''Days before''': Review the [https://etcd.io/docs/latest/upgrades/ etcd upgrade guide] for important points of note relevant to the specific minor version pair you are upgrading from / to given the specific debian releases involved. This is useful for identifying changes to flag default values, etc. that we may need to explicitly set to retain desired behavior.
** For example, if you are reimaging from bookworm to trixie, you would review the 3.4 to 3.5 guide.
** In general, you can disregard the '''Upgrade procedure''' section (if present), as it's not appropriate for our production configuration.
* Depool etcd client traffic from the cluster to be reimaged, as described above.
* Verify that <code>profile::etcd::v3::cluster_bootstrap</code> is false for all conf hosts in the cluster.
** If it is not, you will need to set it to false. Note that applying this change will trigger etcd restarts, so it's preferable to stagger the puppet-agent runs across the 3 cluster member hosts.
{{Note|type=warn|text='''There is no way to depool Zookeeper client traffic''' and clients (e.g., Kafka) will continue to use the cluster while reimages are in flight. This is why it's critical to operate on only one host at a time.}}
==== Reimage ====
One host at a time, execute the following reimage procedure. Do not proceed to the next until both etcd and Zookeeper are healthy.
{{Note|type=info|text=The following primarily makes use of v3 members API <code>etcdctl</code> commands. However, there are some cases where v2 API commands are used, due to complexities around (lack of) gRPC connectivity to all cluster members.}}
* Determine whether etcd-mirror is running on this host (i.e., has <code>profile::etcd::replication::active: true</code>). If so, move it to another cluster member - ideally one that has already been reimaged.
** Prepare a patch that moves <code>profile::etcd::replication::active: true</code> to the host you would like to move to (i.e., defaults to false on the host you are moving from) and update <code>profile::etcd::replication::dst_url</code> to match. Example: {{gerrit|1314016}}.
** [[Alertmanager#Silences_&_acknowledgements|Silence]] the EtcdReplicationDown alert.
** Stop puppet on both hosts: <code>sudo disable-puppet "etcd-mirror move"</code>
** Merge and <code>puppet-merge</code> your patch.
** (replication stops) Run puppet on the host you are moving from (<code>sudo run-puppet-agent -e "etcd-mirror move"</code>). Note that it may take up to 60s for etcd-mirror to terminate.
** (replication starts) Run puppet on the host you are moving to (<code>sudo run-puppet-agent -e "etcd-mirror move"</code>).
* (Optional) Before making destructive changes, if the host you are operating on has not been restarted for some time, consider using the <code>sre.hosts.reboot-single</code> cookbook to verify that you can restart it successfully.
* Determine the member ID of the host you are going to reimage. From any host in the cluster:
<syntaxhighlight lang="bash">
ETCDCTL_API=3 etcdctl --endpoints https://$(hostname -f):2379 member list
</syntaxhighlight>
* Remove the member from the cluster.
** Note that this will cause etcd on the removed member to leave the cluster and terminate ([https://etcd.io/docs/v3.3/op-guide/runtime-configuration/#remove-a-member docs]). You can monitor this process on the host to be removed with, e.g., <code>journalctl -f -u etcd.service</code>.
** From any host in the cluster:
<syntaxhighlight lang="bash">
ETCDCTL_API=3 etcdctl --endpoints https://$(hostname -f):2379 member remove $MEMBER_ID
</syntaxhighlight>
* Start the reimage, e.g., <code>sudo cookbook sre.hosts.reimage --os bookworm -t TXXXXX $MEMBER_HOST_NAME</code> (note: not FQDN).
** The cookbook may recommend you run with <code>--move-vlan</code>. Unless this is what you intend to do, which is a more involved process and not described here, '''do not''' do this.
* Any time after the reimage has started, re-add the host to the cluster. With <code>$MEMBER_FQDN</code> as the member you are reimaging, from any other host in the cluster:
<syntaxhighlight lang="bash">
ETCDCTL_API=3 etcdctl --endpoints https://$(hostname -f):2379 member add $MEMBER_FQDN --peer-urls=https://$MEMBER_FQDN:2380
</syntaxhighlight>
* Wait for the reimage to complete. Upon the first puppet run, <code>etcd.service</code> should start and join the (existing) cluster.
* Confirm that the etcd member has joined the cluster. From any host:
<syntaxhighlight lang="bash">
ETCDCTL_API=3 etcdctl --endpoints https://$(hostname -f):2379 member list
ETCDCTL_API=2 etcdctl -C https://$(hostname -f):2379 cluster-health # v2 API (needs connectivity to all members)
</syntaxhighlight>
* Confirm that the Zookeeper member is healthy and has joined the ensemble.
** From the reimaged host: <code>echo ruok | nc localhost 2181; echo; echo stat | nc localhost 2181</code>.
** From the Zookeeper leader: <code>echo mntr | nc localhost 2181</code> to confirm the number of synced followers (should be 2).
*** Determine the leader by running <code>echo stat | nc localhost 2181</code> on each member, e.g., with cumin (look for <code>Mode: leader</code>).
* Confirm that the nginx etcd TLS proxy is healthy. From any production host, hit 4001 with a known-good path: e.g. <code>curl -v https://$MEMBER_FQDN:4001/v2/keys/</code>.
==== Cleanup ====
* Delete the EtcdReplicationDown silence if it's still active (and not alerting).
* Repool etcd client traffic.
== See also ==
* [[SLO/etcd main cluster]]
3pfjln8djf68uxywlswfu4814t804ke
OAuth
0
443981
2445244
2234801
2026-08-08T21:20:57Z
Tgr (WMF)
39134
update for OAuth 2; add info about edge tokens; add more storage / logging details; remove a bunch of stuff that did not seem important
2445244
wikitext
text/x-wiki
'''OAuth''' is a protocol that lets applications impersonate users in limited ways, after getting permission from those users. It has two versions, OAuth 1.0 (very old and rarely used, but still popular in the Wikimedia tooling comunity) and OAuth 2.0 (widely used on the internet). The public Wikimedia wikis use both versions (provided by [[mw:Extension:OAuth]]), e.g. to allow users to offload editing tasks to external applications (usually, but not always, applications on [[Toolforge]]).
For a description of how OAuth works, and how to develop applications that use it, see [[mw:OAuth/For Developers]]. For a user manual, see [[mw:Help:OAuth]]. This page is focusing on operations concerns. [[mw:OAuth/Owner-only consumers|Owner-only OAuth consumers]] are not discussed here; for the most part those can be thought of as a special kind of password, much like [[mw:Manual:Bot passwords|bot passwords]].
== Edge interactions ==
OAuth 2 identifies requests with an <code>Authorization: Bearer <token></code> header; the token is a JWT signed with the key pair configured in <code>$wgOAuth2PrivateKey</code> / <code>$wgOAuth2PublicKey</code>. ATS and the API gateway validate it and use it for higher rate limits.
OAuth1 identifies requests with an <code>Authorization: OAuth <signature></code> header, but the signature is not decodeable, and does not result in any special behavior in the edge.
== Data ==
The OAuth MediaWiki extension uses three tables: <code>oauth_registered_consumer</code> which stores data about clients (applications using OAuth, more or less; although one application can have several registered clients, mainly as a way of versioning), <code>oauth_accepted_consumer</code> which stores data about the authorizations users give to clients, and <code>oauth2_access_token</code> which stores data about the tokens OAuth 2 uses as credentials. (OAuth 1 credentials are stored in <code>oauth_accepted_consumer.access_token</code> but the process by which they are turned into signatures in the header is more complicated.) See DB schema [[gerrit:plugins/gitiles/mediawiki/extensions/OAuth/+/master/schema/mysql/OAuth.sql|in gitiles]].
The main OAuth installation is on metawiki; clients registered there might be able to access all public Wikimedia wikis (although this might be restricted by the consumer settings). The data tables are also in the metawiki database. Private wikis could in theory have their own OAuth server and their own versions of the tables, but currently none do.
OAuth 2 refresh tokens are stored outside the wiki database, in MainStash ([[phab:T390000|T390000]]).
== Security ==
OAuth applications need to be registered by the developer, and (unless they have very limited rights) go through manual review by [[m:Special:ListUsers/oauthadmin|OAuth admins]] or stewards, which tends to be cursory and mostly just involves checking the permissions the application would have. We almost never approve clients with highly security-sensitive permissions (e.g. Javascript editing), although some exceptions exist. Clients with admin permissions (e.g. page deletion) are generally approved if there is a sound rationale and a trusted community member involved. See [[m:OAuth app guidelines]].
OAuth applications get a client ID / client secret pair (some OAuth 2 apps only use the client ID), which can be exchanged, possibly with user interaction, for an access token (for OAuth 1, an access token / access secret pair) for a specific user. OAuth 1 request signatures require all four values; OAuth 2 only requires the access token. A leaked client secret needs to be reset by the application owner.
See also [https://office.wikimedia.org/wiki/Wikitech/OAuth Wikitech/OAuth on officewiki].
== Management ==
Client management happens via [[m:Special:OAuthManageConsumers]] (you need to be in the <code>oauthadmin</code> group, or steward or staff). Publicly viewable consumer information is at [[m:Special:OAuthListConsumers]]. There is no search functionality but you can follow the link for an arbitrary consumer and replace the client ID in the URL.
In case of a data leak, or an application misbehaving, go to OAuthManageConsumers and set the consumer to ''disabled''. That should prevent any kind of abuse of credentials (reject all API requests made with those credentials), and prevent new users from granting permission to the application. The action can be undone. There's also a "disable and suppress" option in case the registration mechanism itself is abused, e.g. by putting private information in a consumer name.
If a single user is misbehaving and abusing an application, they can simply be blocked on the wiki by name; IP/range blocks don't work though, as the wiki will see the tool IP. Popular tools are expected to provide their own anti-abuse mechanisms - if they don't, and are repeatedly abused, just disable the tool and tell the owner to improve it.
The admin interface only allows disabling/reenabling; most client details can't be changed (the developer needs to create a new consumer). In the case of a leaked client secret, the owner can reset it. Revoking permission from a specific user can only be done by that user, from the "connected applications" link in their preferences. As always, it's best to avoid direct DB manipulation, but it should be reasonably safe if really needed for some reason (short of changing ids/keys).
Many things use OAuth; disabling it on the main cluster is extremely disruptive. If you need to do that anyway, set <syntaxhighlight inline lang=php>$wmgUseOAuth = false</syntaxhighlight> in that wiki/cluster's config. Alternatively, you can set <syntaxhighlight inline lang=php>$wgMWOAuthReadOnly = true</syntaxhighlight> (will disallow any changes to OAuth consumers, but not using existing consumers) or <syntaxhighlight inline lang=php>$wgMWOauthDisabledApiModules[] = "ApiModuleClassName"</syntaxhighlight> (will prevent that API module from being used with any OAuth credentials).
== Logging ==
Consumer state changes (proposal, approval, disabling etc) are logged in the (public) wiki log: [[m:Special:Log/mwoauthconsumer]]. They are also logged to the <code>OAuth</code> channel on Logstash. Authorizations (users giving permission to an application) are also logged to the <code>OAuth</code> channel. Authentication (ie. when an API request uses OAuth credentials) is not logged.
Actions performed via OAuth have an <code>OAuth CID: <id></code> [[mw:Manual:Tags|change tag]] on them. Note this is the consumer ID, not the consumer key. You can find it via the links in [[m:Special:OAuthListConsumers]], or in the database. To get a quick estimate of how important an OAuth tool is, you can look at the <code>change_tag_def.ctd_count</code> table field for that change tag to get the total number of times it was used on the given wiki (there are important tools only doing read actions though).
The OAuth Logstash channel has a [https://logstash.wikimedia.org/app/dashboards#/view/1e16dc40-e07a-11ed-ae74-e713b6900c9a dashboard].
lsxs0irwsib6r9tnoite21qsdmx8pcf
Fundraising/Data and flow/Audits
0
448785
2445249
2443738
2026-08-09T03:46:29Z
Quiddity
1884
lang="bash"
2445249
wikitext
text/x-wiki
= Reconciliation (Audit framework) =
We have a framework to reconcile the donations/refunds/chargebacks etc in CiviCRM with the reports we receive from the payment processors.
We do this to achieve 2 goals
1) ensure that all donor gifts are shown in CiviCRM (sometimes other notifications are not received) and that all the information we want is stored. Some information, like the fees, the converted settled amount and the reference for the batch the contribution settles in first becomes available to us when we receive audit/reconcilation reports from the processors. (
2) ensure that we can accurately send finance batches of the amouts settled by each of our settlement providers (currently adyen, paypal, braintree, dlocal, trustly, checkout.com and chariot). There is further data on this on the [[Fundraising/Data and flow/Intacct|Fundraising/Data and flow/Intacct page]] - note that that page is intended to be readable by people outside fr-tech
The priorities are slightly different for the 2 goals. For the first goal our priority is to get the donations in as soon as we can - so we process 'whatever we get, when we get it'. In practice this means we process both payment and settlement reports from the payment processors and payment reports from Gravy. (In most cases the first report we receive is the processor payment report and the other 2 are for redundancy). We only move the report from 'incoming' to 'completed' when all transactions in it are in CiviCRM.
For the second goal our priority is to verify the exact total that settled in a batch and ensure that all the transactions that contributed to that total are recorded in CiviCRM against the batch with final amount information in the settled currency. Determining the exact total varies by payment processor so the table below shows how we reach this amount based on the information we receive from each payment processor.
The information from the settlement providers is our primary source of information. Gravy also receives these reports in some cases and returns them to us with minor formatting changes. Gr4vy sales people have promoted their reports as a better source but in practice we have found that they often don't exist until months after we have implemented a processor, and they are inferior to those from the payment providers as they don't provide settlement batch total information for validation purposes, and they don't always match what is settled, especially around the more obscure scenarios. We do the additional work to get the gravy ones working but are not quite sure if we gain anything through this so we focus our efforts on the reports we get from the primary source. We cannot use gravy reports instead of our process as all records must be matched with CiviCRM donations in order to enhance them with GL-related information.
We retain the information we get from the primary sources on disk and can, if necessary, access it from there. We currently never delete them and have primary data from PayPal and Adyen going back over 12 years!
{| class="wikitable"
|+
!
!Report type
!Name
!Batch total calculation
!Gravy report availability
!Schedule, files, notes
|-
|Adyen
[[Fundraising/Data and flow/PSP integrations/Adyen Checkout#Audits]]
|csv Settlement report
|*settlement_batch*
|Adyen has a row `payout` and each transaction has a batch number. The records with the batch number add up to the batch payout row.
|Available, mostly identical to what we get from Adyen directly (column names differ) but not cover obscure adjustments in all cases
|We receive 3 reports - 2 versions of the settlement report with different names and one payment report.
- payments_account_report - nightly file, has the previous days transactions. comes in a little sooner so by processing it as well we are able to start the processing a little earlier.
- settlement_detail_report_batch - daily now that we settle daily
settlement_detail_xxxx_batch - same weekly file with a different name - we get both because Gravy wanted a different name. Currently parsing both due to perceived risk that Gravy requirements will change and one will go away & it will be the one we are parsing.
Runs multiple times a day but we only parse one file at a time to try to avoid double queueing to the settle queue - less of an issue now we skip the gravy-requested file
|-
|PayPal
|csv Settlement report
|STL-***
|PayPal has a footer section that gives the details for amounts by currency. We calculate the file total as Credits - Debits + Fee Credits - Fee Debits . To get the batch total we then deduct any Debits that relate to accounts transfers or expense re-imbursements
|Available with Data gaps:
No information to allow us to determine batch total.
Some transactions missing - eg. https://phabricator.wikimedia.org/T418191 affects 4 settlement batches but is only present in 2 of the gravy files (skipped)
Otherwise mostly the same as the primary source (some formatting changes)
|We only process the TTR & STL reports, Others are moved to ignored. TTR is the payments report.
Runs nightly at 17:35 UTC
TRR - nightly Transaction Detail Report - this is used by the audit and has all the single transactions
STL - Settlement Report - this is used by the audit and has all the settled transactions
SAR - nightly Subscription Agreement Report - this 'disappeared' - possibly Gravy related.
WIkimedia - not used
|-
|Trustly
|csv Settlement rport
|P11KFUN-
|Trustly provides the batch total in the footer and in payout rows. We use the payout rows (in case there is ever more than one per file)
|n/a - they did announce it was 'there' manyl months after we were up & running with reports direct from Trustly but I did check on 21 Jul 2026 as part of this documenation update & don't see them
|
|-
|Dlocal
(formerly Astropay)
[[Fundraising/Data and flow/PSP integrations/dLocal#Audits]]
|csv Settlement report
|cross_border
|Dlocal provides the batch total in the header. The calculations are all done before rounding (5 decimal places) so we calculate a rounding transaction which we code as a fee
|n/a
|Runs every night at 00:20 UTC
files - nightly
|-
|Braintree
|api - returned in json
Using graphQL
|
|We retrieve the transactions by disbursement date. We have to add these up to get the batch total. I went through a period of manually verifying these against the Disbursement reports in the UI and these calculated totals accurately reflect the Donations + Refunds that show up in the UI.
Chargebacks sit outside those 2 - they are not counted in the main disbursement reports in the UI or retrievable by disbursement dates. These are low volume and we are currently treating them as separate batches.
|n/a
|Runs every night at 00:00 UTC
2 json files created - one for transactions and refunds and the other for chargebacks (if any). Chargebacks are a bit messy in that we can't search by the date the chargeback settled so we try using a wider window to try and catch them - seems to work
|-
|Stripe
|csv
|payout row
|
|n/a
|
|-
|CheckoutCom
|csv settlement report with a payout report to specify payout amount
|payouts + settlement_breakdown
|We get the transactions from the settlement breakdown and the final payout from the payouts report. We round if there is a misalignment
|n/a
|
|-
|Chariot
|api - returned in json
|deposits + donations
|When a deposit is available we retrieve the donations for it
|n/a
|
|}
====== '''Gravy parsing note''' ======
Runs every night at 02:10 UTC
files - nightly at 01:00 UTC. We only parse the payments file and this appears to run after other files so is a bit of a null op.
=== Process control jobs for the audit parsing ===
This documents the workflow to process audit files from payment processors and import missing messages into CiviCRM.
Most payment processors have 1 or 2 audit jobs that we run multiple times a day (some only once a day). We follow the job naming pattern of *_audit_download and *_audit_parse - eg.
braintree_audit_download
braintree_audit_parse
The download jobs generally run a custom Maintenance script located in the smash-pig standalone codebase that will either retrieve and compile data from the processor's api or download files from an SFTP location. For the latter there is a generic script. When compiling from the api we store either as raw json or converted to a csv depending on the data source complexity (e.g braintree is stored as a json as that can be parsed fairly easily from the raw json but for chariot we are combining data from 2 api calls so the `GetReport` code combines these to a csv with a lot of processing.
The WMFAudit.parse api call processes the reconciliation files. These files will contain an api call like
<code>wmf-cv api4 -vv WMFAudit.parse gateway=braintree</code>
The code in the civicrm extension handles reading the list of files from the directory, searching for existing transactions in the database, and finding missing information in the smashpig.pending tables (with a legacy fall back to searching payments-wiki logs (mounted at /srv/archive/frlog/logs)) for each transaction that isn't in the database. The code to parse the individual files to an array of normalized transactions lives under the SmashPig codebase, in classes that implement the [[phab:diffusion/WFSP/browse/master/Core/DataFiles/AuditParser.php|AuditParser interface]], whereas the code to reconcile those files lives in the CiviCRM wmf-civicrm extension.
Audit files are located on civi1001 in <code>/var/spool/audit/[payment-processor]</code> and divided into two directories: incoming and completed.
==== How to run the parser (works locally too) ====
Run the audit parser for one gateway:
<code>wmf-cv api4 -vv WMFAudit.parse gateway=adyen logSearchPastDays=12</code>
Run with just one specific file:
<code>wmf-cv api4 -vv WMFAudit.parse gateway=adyen logSearchPastDays=12 file=payments_accounting_report_2024_04_06.csv</code>
====== Additional optional parameters ======
- settleMode ("queue" or "now", blank is implicitly false) - should the settle queue be populated from these
- file (specify file name) - if set then only one file will be processed
- isStopOnFirstMissing (bool) - primarily for debug usage
- rowLimit (int) - primarily for debug usage
- offset (int) -primarily for debug usage
- l''ogInterval'' ''(int) - how often should progress be output''
- isMoveCompletedFile (bool) should the file be moved afterwards - mostly for test & debug
- isCompleted (bool) look in the completed folder - test & debug usage
==== <b>Refunds & Donations getting into Civi & getting settled</b> ====
Missing donations will be queued to CiviCRM. If they are in CiviCRM but the audit has additional settlement data that will be added to the settlement queue - these queues can be processed using process-control or with the api
{| class="wikitable"
|+
!Process control
!API
|-
|run-job -j donations_queue_consume
|<code>wmf-cv api4 -vv WMFQueue.Consume timeLimit=280</code> queueConsumer=Donation queueName=donations
|-
|run-job -j refund_queue_consume
|<code>wmf-cv api4 -vv WMFQueue.Consume timeLimit=280 queueConsumer=Refund queueName=refund</code>
|-
|run-job -j settle_queue_consume
|wmf-cv api4 -vv WMFQueue.Consume timeLimit=173 queueConsumer=Settle queueName=settle
|}
=== Resolving Audit issues ===
Audit issues come to fr-tech attention in one of 2 ways
# fail mail
# the daily finance batch summary email
In the case of fail mail the job has failed and you should either look at the log output (see below) or try re-running the job & look at the output you get.
====== For the emails there are a few things to check ======
1) if the subject line says ''''1 batch needs attention'''` this means that there is a NEW issue with a batch not adding up to the total it should settle to (or at least a new batch with an issue, if previously identified). The batch will be updated at this point to having a status of 'Needs attention' and will not trigger that subject line on subsequent days.
2) if the subject line mentions ''''files older than'''<nowiki/>' this means some files are failing to clear - this likely means some transactions in the audit files are not being matched to donations in CiviCRM or have failed to create donations
3) if the subject line says one '''contribution needs attention''' it means it has not identified a GL code for this contribution. As of Aug 2026 these are almost definitely offline contributions that have not been tagged with 'is_major_gift' - and the quick answer is to add is_major_gift to that specific donation - we were still finalising the definition of is_major_gift in Prague and have not revisited since. If this gets painful we could do an interim code fix but it's probably going to be looked at again properly imminently.
4) Other than the subject line there is a table further down of open & needs attention batches. These should be checked to make sure what is there makes sense. Any that are 'needs attention' should have a phab next to them and should be being actively worked on. Any that are 'a bit old' may also need some attention
==== How to resolve issues ====
- if the file has not cleared and / or you have a fail mail then either try re-running the job or check the log output per the log output section below
- if there is an issue with a batch then you will need to figure out why it is not balancing. Once you have resolved it you will need to reset the status (generally I reset it to 'open' and then re-run the audit file allowing that to update it to `total_verified` - but ultimately it needs a status of total_verified or validated before it can be exported. It can be hard to track down why these don't balance but here are some things that might help
# check the [https://frmon.frdev.wikimedia.org/d/Pq1YNMviz/fundraising-overview?orgId=1&from=now-24h&to=now&timezone=utc&var-payments_host=payments1005&var-payments_host=payments1006&var-payments_host=payments1007&refresh=1m&viewPanel=panel-11 settle queue is clear] - in some cases it just ran before all settlement transactions were processed
# you can check all batch statuses at https://civicrm.wikimedia.org/civicrm/accounting/batches and link through from there to see the transactions
# You can try re-validating the batch to check / get more information - you can do that in the ui screen above but usually the [https://docs.civicrm.org/dev/en/latest/api/v4/usage/#cv CLI] is better - eg in this case I can see the credits (donations) match but there was an expected debit of -3.10 which is missing - I now know I'm looking for negative transaction of $3.10 that did not get to CiviCRM<syntaxhighlight lang="bash">
wmf-cv -vv api4 Batch.validate +w id=8535
"totals": {
"debit": "0.00",
"credit": "298.86",
"fee_debit": "0.16",
"fee_credit": "0.00",
"fee": "0.16",
"settled": "298.70",
"count": 42
},
"expected": {
"count": 43,
"credit": 298.86,
"debit": -3.1,
"fee": -0.16,
"settled": 295.6
},
"validation": {
"count": 1,
"credit": 0,
"debit": -3.1,
"fee": 0,
"settled": -3.1
}
</syntaxhighlight>
# Generally you will need to access the source file and try to find what is in there but not in CiviCRM. The view that is linked in the email and from the batch summary ([https://civicrm.wikimedia.org/civicrm/contribution/settled#?finance_batch=trustly_3442702_USD e.g]) gives you scope to filter on donations of various amounts to try to bisect what is missing. You will find the file in the folder in the setting `wmf_audit_directory_audit` - eg. try <syntaxhighlight lang="bash">
wmf-cv Setting.get | grep audit
</syntaxhighlight>within that directory you are looking in either (eg) trustly/incoming or trustly/completed e.g I have found that one of the transactions in the batch I'm looking for has the transaction ID 8197772279 so from the trustly directory: <syntaxhighlight lang="bash">
grep 8197772279 */*
zgrep 8197772279 */*
</syntaxhighlight>Once you have found the file you might find it easier to move it back to incoming & unzip it. Re-running the parse on it *might* give a clue
#In this case I found a row in the csv grepping for 3.10 with a status of 'reversed' and then looked to see if it was in Civi & on not finding it I felt ready to log a phab - after logging the phab go back to the batch summary and update the batch to have the phab next to it.
====== Log output ======
check audit parse log from frlog1002, located at /srv/archive/civi/process-control/<yyyymmdd>/<paymentMethod>_audit_parse-<yyyymmdd>-xxxxxx.log.civi1001.bz2 - you can also just re-run the job at any time & see what it does.
Generally the logs will show you the first transaction that was missing from CiviCRM and if this transaction persists after the file has been processed once then you should generally investigate and resolve this transaction and then try re-running it.
Result print example as:
Done! Final stats:
Total number of donations in audit file: xxx
Number missing from database: xxx
Missing transactions found in logs: xxx
Missing transactions not found in logs: xxx
Missing transaction summary:
Regular donations: 2
Returned from hook drush_wmf_audit_parse_audit [1.13 sec, 36.01 MB] [debug]
xxxxxxxx: 1
xxxxxxxx: 1
Refunds and chargebacks: 0
Recurring donations: 0
Command dispatch complete [1.13 sec, 35.92 MB] [notice]
Transaction IDs:
xxx xxxxxxxxxx
xxx xxxxxxxxxx
Initial stats on recon files: Array
(
[/var/spool/audit/xxx/incoming/xxx] => 0
[/var/spool/audit/xxx/incoming/xxx] => 0
[/var/spool/audit/xxx/incoming/xxx] => 2
)
== File Wrangling ==
Sometimes the audit processor can't resolve all the transactions in a file, even after trying for several days. This can lead to a build-up of files in the incoming directory and to subsequent processor runs getting longer and longer till finally they start timing out. Generally it is best to stay on top of these and to run individual files if needed. If files are not processed they will not get to Intacct.
As a last resort we can manually temporarily move the older files from the incoming to the completed directory. Since our personal accounts don't have permissions to move the files, we do this with a one-off process-control job such as ingenico_move_audit_files. Since process control runs each command as a separate process under python, we need to wrap any file globs that we want expanded with 'sh -c', for example:
<pre>sh -c "mv /srv/archive/civi1001/audit/globalcollect/incoming//wx1*202010[01][0-5]*xml* /var/spool/audit/globalcollect/"</pre>
== Adding a New Payment Processor ==
Once the new processor code has been added
1. Enable the Module
2. Add the folders
This is done via puppet by adding the processor name to the <code>$audit_processors</code> array in <code>modules/civicrm/manifests/audit.pp</code>. Additionally, some files (YAML usually) may be needed for the audit configuration. Those vary by processor but are stored in the same <code>audit.pp</code> manifest.
=== Legacy processors .... ===
==== Amazon ====
Instead of an SFTP download, we have to call methods on the Amazon Pay SDK to get our reports. This is kicked off with the [https://phabricator.wikimedia.org/diffusion/WFSP/browse/master/PaymentProviders/Amazon/Audit/DownloadReports.php DownloadReports] php script in SmashPig.
==== Fundraiseup ====
Process control jobs
fundraise-up_audit.yaml - Runs at 1AM UTC daily. Calls fundraise-up_audit_download and then fundraise-up_audit_parse.
fundraise-up_audit_download - Downloads fundraiseup export files from the Fundraiseup SFTP server. This files contains the new donations, new recurrings, cancelled recurring, and refunded transactions from Fundraiseup.
fundraise-up_audit_parse - Imports the transactions from the exported files. Calls <code>cv api4 --user=admin -vv WMFAudit.parse gateway=fundraiseup logSearchPastDays=12</code>
e761h20xfyujnceygf6v11qvnsls5fm
Map of database maintenance
0
449160
2445245
2445226
2026-08-09T00:02:36Z
Dexbot
30554
Bot: Updating the report
2445245
wikitext
text/x-wiki
{{/Header}}
== Today (2026-08-09) ==
== Yesterday (2026-08-08) ==
== Last seven days ==
{| class="wikitable"
|+ eqiad
|-
! Section !! Work
|-
| x1 || [[phab:T433990|Optimize echo tables in x1 (T433990)]] (ladsgroup)
|-
|}
{| class="wikitable"
|+ codfw
|-
! Section !! Work
|-
| s4 || [[phab:T433610|FIRING: [2x] SystemdUnitFailed: ifup@eno12399np0.service on db2248:9100 (T433610)]] (marostegui)
|-
| x1 || [[phab:T433990|Optimize echo tables in x1 (T433990)]] (ladsgroup)
|-
|}
[[Category:MariaDB]]
3wwaixq2njdvwwxkbgx6wdyfgkr2dq4
Tool:Gitlab-account-approval/Log
116
453906
2445251
2445243
2026-08-09T06:42:15Z
Gitlabaccountapprovalbot
37332
marsam2489 was rejected.
2445251
wikitext
text/x-wiki
<noinclude>'''Audit log of approvals''' made by [[gitlab:gitlabaccountapprovalbot|@gitlabaccountapprovalbot]]. __NOTOC__</noinclude>
=== 2026-08-09 ===
* 06:42 "marsam2489" was rejected (pending since 2026-05-10T06:40:46.276Z).
=== 2026-08-08 ===
* 08:15 [[gitlab:taiwaniajusto|@taiwaniajusto]] was approved.
=== 2026-08-07 ===
* 18:21 [[gitlab:gturkington|@gturkington]] was approved.
* 06:36 "brianbybyby" was rejected (pending since 2026-05-08T06:36:08.459Z).
=== 2026-07-30 ===
* 20:42 "horaciocolbert" was rejected (pending since 2026-04-30T20:41:40.421Z).
=== 2026-07-29 ===
* 10:03 [[gitlab:piastu|@piastu]] was approved.
* 06:33 "rafiul1" was rejected (pending since 2026-04-29T06:33:02.573Z).
=== 2026-07-28 ===
* 16:51 [[gitlab:for-each-next|@for-each-next]] was approved.
* 09:21 [[gitlab:gka|@gka]] was approved.
=== 2026-07-27 ===
* 08:12 [[gitlab:cambob|@cambob]] was approved.
=== 2026-07-26 ===
* 15:39 "demansanaagmailcom" was rejected (pending since 2026-04-26T15:39:03.297Z).
=== 2026-07-24 ===
* 10:54 [[gitlab:ysogo|@ysogo]] was approved.
* 09:39 [[gitlab:pankaj199|@pankaj199]] was approved.
=== 2026-07-23 ===
* 12:30 [[gitlab:fermiboson|@fermiboson]] was approved.
* 09:21 [[gitlab:slashme|@slashme]] was approved.
=== 2026-07-22 ===
* 15:54 [[gitlab:lmedley|@lmedley]] was approved.
* 13:24 "praveen5638" was rejected (pending since 2026-04-22T13:21:23.368Z).
* 12:48 [[gitlab:cyberpower678|@cyberpower678]] was approved.
* 08:30 [[gitlab:nabbegat|@nabbegat]] was approved.
* 08:30 [[gitlab:plyd|@plyd]] was approved.
* 07:06 "ayush8620" was rejected (pending since 2026-04-22T07:03:12.476Z).
=== 2026-07-21 ===
* 15:51 [[gitlab:panieravide|@panieravide]] was approved.
* 15:33 [[gitlab:luisvilla-personal|@luisvilla-personal]] was approved.
* 13:48 [[gitlab:yru|@yru]] was approved.
* 13:09 [[gitlab:deevad|@deevad]] was approved.
* 12:54 [[gitlab:ctdo17|@ctdo17]] was approved.
* 12:36 [[gitlab:jeannenoiraud|@jeannenoiraud]] was approved.
* 12:21 [[gitlab:nadiantara|@nadiantara]] was approved.
* 12:18 [[gitlab:wijltcher|@wijltcher]] was approved.
* 10:33 [[gitlab:nivopol|@nivopol]] was approved.
* 10:30 [[gitlab:johlig|@johlig]] was approved.
* 10:30 [[gitlab:majicita|@majicita]] was approved.
* 10:15 [[gitlab:yongjiapeng|@yongjiapeng]] was approved.
* 10:09 [[gitlab:francyskus|@francyskus]] was approved.
* 09:24 [[gitlab:xanonymusx|@xanonymusx]] was approved.
=== 2026-07-20 ===
* 17:15 "leonidlednev" was rejected (pending since 2026-04-20T17:13:35.108Z).
* 15:27 [[gitlab:rodrigoargenton|@rodrigoargenton]] was approved.
* 05:48 "draftecho" was rejected (pending since 2026-04-20T05:48:06.953Z).
=== 2026-07-19 ===
* 12:21 [[gitlab:boivie|@boivie]] was approved.
=== 2026-07-18 ===
* 16:09 [[gitlab:pharos|@pharos]] was approved.
* 15:45 [[gitlab:priyankar22|@priyankar22]] was approved.
* 15:30 [[gitlab:sisyph|@sisyph]] was approved.
=== 2026-07-13 ===
* 03:45 [[gitlab:dreamyshade|@dreamyshade]] was approved.
=== 2026-07-12 ===
* 09:27 [[gitlab:smk|@smk]] was approved.
=== 2026-07-11 ===
* 14:48 "bigcereal42" was rejected (pending since 2026-04-11T14:47:40.321Z).
* 12:18 "pratyushsawan" was rejected (pending since 2026-04-11T12:16:41.671Z).
=== 2026-07-10 ===
* 14:42 [[gitlab:akaza24|@akaza24]] was approved.
* 12:57 [[gitlab:kormisk|@kormisk]] was approved.
=== 2026-07-07 ===
* 10:36 [[gitlab:olafjanssen|@olafjanssen]] was approved.
* 06:57 "elisapoly-99" was rejected (pending since 2026-04-07T06:55:52.662Z).
=== 2026-07-06 ===
* 11:57 "ma3rouf" was rejected (pending since 2026-04-06T11:56:30.978Z).
=== 2026-07-02 ===
* 15:48 [[gitlab:tekneos|@tekneos]] was approved.
=== 2026-07-01 ===
* 15:12 [[gitlab:mugurolevy|@mugurolevy]] was approved.
* 14:15 [[gitlab:vadymts1|@vadymts1]] was approved.
* 09:57 "mugurolevy" was rejected (pending since 2026-04-01T09:55:19.175Z).
=== 2026-06-30 ===
* 14:27 "shivangisharma" was rejected (pending since 2026-03-31T14:26:44.932Z).
=== 2026-06-29 ===
* 19:03 [[gitlab:thisismattmiller|@thisismattmiller]] was approved.
=== 2026-06-28 ===
* 14:51 "nkwenuinadine" was rejected (pending since 2026-03-29T14:48:32.735Z).
* 14:03 "vaishnavikumbhar" was rejected (pending since 2026-03-29T14:01:30.604Z).
* 13:03 "stepmay" was rejected (pending since 2026-03-29T13:01:59.905Z).
* 06:51 "swallroth" was rejected (pending since 2026-03-29T06:49:54.838Z).
=== 2026-06-26 ===
* 09:39 [[gitlab:lakshita28|@lakshita28]] was approved.
* 07:30 [[gitlab:reeti|@reeti]] was approved.
* 07:30 [[gitlab:anushka10patel|@anushka10patel]] was approved.
* 07:30 "samsaesque" was rejected (pending since 2026-03-27T07:29:57.279Z).
* 05:51 [[gitlab:arpithhhaaa|@arpithhhaaa]] was approved.
* 05:51 [[gitlab:govindlaltl|@govindlaltl]] was approved.
=== 2026-06-25 ===
* 16:09 [[gitlab:sakuraemad|@sakuraemad]] was approved.
* 07:00 "kdh8219" was rejected (pending since 2026-03-26T06:58:05.415Z).
=== 2026-06-24 ===
* 11:54 [[gitlab:sanskardubeydev|@sanskardubeydev]] was approved.
* 10:09 "tanmay789q" was rejected (pending since 2026-03-25T10:07:54.602Z).
=== 2026-06-22 ===
* 19:57 [[gitlab:gouvernathor|@gouvernathor]] was approved.
* 16:45 [[gitlab:lucasbelo|@lucasbelo]] was approved.
* 07:15 "jason2000-cpu" was rejected (pending since 2026-03-23T07:14:09.184Z).
=== 2026-06-21 ===
* 13:18 [[gitlab:egonw|@egonw]] was approved.
=== 2026-06-20 ===
* 10:21 [[gitlab:tways2017|@tways2017]] was approved.
=== 2026-06-19 ===
* 16:06 "wilsonwang2026" was rejected (pending since 2026-03-20T16:06:05.511Z).
* 04:12 [[gitlab:claudio|@claudio]] was approved.
=== 2026-06-18 ===
* 14:21 "royiswariii" was rejected (pending since 2026-03-19T14:19:16.896Z).
* 13:06 [[gitlab:laurabarluzzi|@laurabarluzzi]] was approved.
=== 2026-06-17 ===
* 11:24 "adinathq8x" was rejected (pending since 2026-03-18T11:22:50.098Z).
* 09:45 "nathanveritas" was rejected (pending since 2026-03-18T09:43:51.645Z).
=== 2026-06-15 ===
* 22:39 [[gitlab:mohammadhijjawi|@mohammadhijjawi]] was approved.
* 14:24 "enlisar" was rejected (pending since 2026-03-16T14:23:00.109Z).
* 14:06 "ayaan" was rejected (pending since 2026-03-16T14:03:31.071Z).
* 10:54 "kwametech" was rejected (pending since 2026-03-16T10:54:11.083Z).
=== 2026-06-14 ===
* 17:45 [[gitlab:surajseth520|@surajseth520]] was approved.
* 07:24 "malahimhaseeb" was rejected (pending since 2026-03-15T07:21:57.748Z).
=== 2026-06-11 ===
* 11:48 [[gitlab:cadddr|@cadddr]] was approved.
* 11:18 "wikipiggy" was rejected (pending since 2026-03-12T11:16:09.335Z).
* 07:15 [[gitlab:vesihiisi|@vesihiisi]] was approved.
=== 2026-06-10 ===
* 07:03 [[gitlab:dmiranda|@dmiranda]] was approved.
=== 2026-06-09 ===
* 14:21 [[gitlab:linkgenetic|@linkgenetic]] was approved.
* 14:03 [[gitlab:sjones-ctr|@sjones-ctr]] was approved.
* 12:51 [[gitlab:ekrem|@ekrem]] was approved.
=== 2026-06-08 ===
* 17:48 "jmprax" was rejected (pending since 2026-03-09T17:46:38.807Z).
* 16:15 [[gitlab:ahonc|@ahonc]] was approved.
* 12:57 [[gitlab:rainmonger|@rainmonger]] was approved.
=== 2026-06-07 ===
* 23:03 "shadowthewuff" was rejected (pending since 2026-03-08T23:00:53.442Z).
* 11:45 "wiki-pavan" was rejected (pending since 2026-03-08T11:45:11.116Z).
* 02:39 [[gitlab:launchpad|@launchpad]] was approved.
=== 2026-06-06 ===
* 14:54 "unicord" was rejected (pending since 2026-03-07T14:52:04.992Z).
* 12:48 "chien" was rejected (pending since 2026-03-07T12:48:11.669Z).
=== 2026-06-04 ===
* 14:33 "only-vikas" was rejected (pending since 2026-03-05T14:32:09.186Z).
=== 2026-06-03 ===
* 15:00 [[gitlab:anafibnshahibul|@anafibnshahibul]] was approved.
=== 2026-06-02 ===
* 21:21 "mgagat" was rejected (pending since 2026-03-03T21:18:37.223Z).
* 13:57 "prasunaenumarthy" was rejected (pending since 2026-03-03T13:57:14.847Z).
* 05:48 [[gitlab:tmoney|@tmoney]] was approved.
=== 2026-06-01 ===
* 14:57 "vikram2101" was rejected (pending since 2026-03-02T14:54:26.550Z).
* 12:03 "watshell" was rejected (pending since 2026-03-02T12:03:09.329Z).
=== 2026-05-29 ===
* 12:48 "mounikapotladurthi" was rejected (pending since 2026-02-27T12:45:38.609Z).
=== 2026-05-27 ===
* 20:00 "vinitha" was rejected (pending since 2026-02-25T19:58:43.524Z).
* 16:30 "codeurluce" was rejected (pending since 2026-02-25T16:28:53.973Z).
* 14:33 [[gitlab:thilio|@thilio]] was approved.
=== 2026-05-26 ===
* 12:09 "charisad" was rejected (pending since 2026-02-24T12:07:21.881Z).
=== 2026-05-25 ===
* 22:54 "ddshelto" was rejected (pending since 2026-02-23T22:52:44.427Z).
* 19:51 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z).
* 19:48 "lakz-99" was rejected (pending since 2026-02-23T19:47:00.263Z).
=== 2026-05-24 ===
* 18:45 "jiyagupta-cs" was rejected (pending since 2026-02-22T18:43:33.176Z).
=== 2026-05-23 ===
* 13:09 [[gitlab:gauthammohanraj|@gauthammohanraj]] was approved.
* 04:21 [[gitlab:staraction|@staraction]] was approved.
=== 2026-05-22 ===
* 19:03 "i-horich" was rejected (pending since 2026-02-20T19:00:43.519Z).
* 01:48 "50323233" was rejected (pending since 2026-02-20T01:48:05.555Z).
=== 2026-05-21 ===
* 18:51 "kartikeyg0104" was rejected (pending since 2026-02-19T18:48:39.707Z).
* 16:27 [[gitlab:renovatebot|@renovatebot]] was approved.
* 16:06 [[gitlab:gkm563|@gkm563]] was approved.
=== 2026-05-20 ===
* 01:21 "beedellrokejulianlockhart" was rejected (pending since 2026-02-18T01:19:13.284Z).
=== 2026-05-18 ===
* 23:18 "wladek92" was rejected (pending since 2026-02-16T23:16:22.939Z).
* 16:36 [[gitlab:effeietsanders|@effeietsanders]] was approved.
=== 2026-05-14 ===
* 21:00 [[gitlab:nehemienathan|@nehemienathan]] was approved.
=== 2026-05-13 ===
* 10:51 "ssssaaaa" was rejected (pending since 2026-02-11T10:50:36.975Z).
=== 2026-05-12 ===
* 18:06 [[gitlab:psubhashish|@psubhashish]] was approved.
* 08:12 "khan" was rejected (pending since 2026-02-10T08:11:48.776Z).
* 04:27 "galaxysh" was rejected (pending since 2026-02-10T04:24:59.440Z).
=== 2026-05-11 ===
* 12:18 "peterxy12" was rejected (pending since 2026-02-09T12:18:01.982Z).
=== 2026-05-10 ===
* 11:09 "yalihupokn" was rejected (pending since 2026-02-08T11:06:51.336Z).
* 05:12 "wobadha" was rejected (pending since 2026-02-08T05:11:00.569Z).
=== 2026-05-09 ===
* 13:45 "bwiki" was rejected (pending since 2026-02-07T13:43:38.177Z).
=== 2026-05-08 ===
* 09:24 [[gitlab:cwilliams|@cwilliams]] was approved.
=== 2026-05-07 ===
* 14:15 "rehankhan78" was rejected (pending since 2026-02-05T14:13:37.754Z).
=== 2026-05-06 ===
* 11:24 "ari" was rejected (pending since 2026-02-04T11:24:11.760Z).
* 08:09 [[gitlab:neriah|@neriah]] was approved.
* 06:27 [[gitlab:status401|@status401]] was approved.
=== 2026-05-03 ===
* 09:54 [[gitlab:anilk|@anilk]] was approved.
=== 2026-05-02 ===
* 17:54 [[gitlab:sweil|@sweil]] was approved.
* 17:00 [[gitlab:aoppo|@aoppo]] was approved.
=== 2026-05-01 ===
* 21:18 [[gitlab:dawalda|@dawalda]] was approved.
=== 2026-04-30 ===
* 21:42 "merohibine" was rejected (pending since 2026-01-29T21:40:00.756Z).
* 20:54 [[gitlab:tfmorris|@tfmorris]] was approved.
* 17:33 [[gitlab:uyen|@uyen]] was approved.
* 07:39 [[gitlab:mahveotm|@mahveotm]] was approved.
* 06:36 [[gitlab:leo321|@leo321]] was approved.
=== 2026-04-29 ===
* 02:27 [[gitlab:dw31415|@dw31415]] was approved.
=== 2026-04-28 ===
* 23:09 [[gitlab:dtorsani|@dtorsani]] was approved.
=== 2026-04-27 ===
* 23:42 [[gitlab:quinlan|@quinlan]] was approved.
* 05:00 [[gitlab:matthewyeager|@matthewyeager]] was approved.
=== 2026-04-26 ===
* 17:36 "kuba-hajnej" was rejected (pending since 2026-01-25T17:33:32.467Z).
* 13:03 "jklamo" was rejected (pending since 2026-01-25T13:02:22.936Z).
=== 2026-04-25 ===
* 20:24 [[gitlab:maldaxura|@maldaxura]] was approved.
* 14:33 [[gitlab:sirtobi|@sirtobi]] was approved.
* 04:18 "ice5678" was rejected (pending since 2026-01-24T04:15:30.008Z).
=== 2026-04-24 ===
* 22:06 [[gitlab:arcstur|@arcstur]] was approved.
=== 2026-04-22 ===
* 23:06 "dtorsani" was rejected (pending since 2026-01-21T23:03:25.843Z).
* 22:18 [[gitlab:egezort|@egezort]] was approved.
* 16:45 "nexpectarpit" was rejected (pending since 2026-01-21T16:43:21.045Z).
=== 2026-04-20 ===
* 19:15 "fitch" was rejected (pending since 2026-01-19T19:12:35.644Z).
=== 2026-04-19 ===
* 02:54 [[gitlab:neoact|@neoact]] was approved.
=== 2026-04-18 ===
* 07:06 [[gitlab:kockaadmiralac|@kockaadmiralac]] was approved.
=== 2026-04-17 ===
* 13:42 "liselot" was rejected (pending since 2026-01-16T13:39:41.909Z).
=== 2026-04-15 ===
* 17:03 "lahari" was rejected (pending since 2026-01-14T17:02:06.275Z).
=== 2026-04-14 ===
* 13:00 "surajseth520" was rejected (pending since 2026-01-13T12:59:45.906Z).
* 04:51 [[gitlab:canley|@canley]] was approved.
* 01:03 "bshizzle" was rejected (pending since 2026-01-13T01:00:48.120Z).
=== 2026-04-13 ===
* 15:30 [[gitlab:passimacopoulos|@passimacopoulos]] was approved.
=== 2026-04-11 ===
* 12:30 "krithash" was rejected (pending since 2026-01-10T12:27:24.731Z).
=== 2026-04-10 ===
* 15:30 "raunak1709" was rejected (pending since 2026-01-09T15:29:10.901Z).
=== 2026-04-07 ===
* 17:03 [[gitlab:supnabla|@supnabla]] was approved.
=== 2026-04-06 ===
* 20:00 [[gitlab:laerdon|@laerdon]] was approved.
* 19:21 [[gitlab:ljq3|@ljq3]] was approved.
=== 2026-04-04 ===
* 11:06 "mixcc" was rejected (pending since 2026-01-03T11:03:33.922Z).
=== 2026-04-02 ===
* 05:30 [[gitlab:mbh1|@mbh1]] was approved.
=== 2026-04-01 ===
* 18:21 "yuvrajpatil17" was rejected (pending since 2025-12-31T18:20:27.991Z).
* 12:12 [[gitlab:amorii0|@amorii0]] was approved.
=== 2026-03-31 ===
* 11:00 "krrishsehgal" was rejected (pending since 2025-12-30T11:00:16.384Z).
=== 2026-03-30 ===
* 15:36 [[gitlab:atsuko|@atsuko]] was approved.
=== 2026-03-29 ===
* 11:36 [[gitlab:giftcup|@giftcup]] was approved.
=== 2026-03-28 ===
* 14:51 [[gitlab:janeeva1|@janeeva1]] was approved.
=== 2026-03-26 ===
* 13:36 [[gitlab:saiphani02|@saiphani02]] was approved.
* 11:48 [[gitlab:valerioboz-wmch|@valerioboz-wmch]] was approved.
=== 2026-03-25 ===
* 09:45 "quansi" was rejected (pending since 2025-12-24T09:42:13.451Z).
* 02:18 [[gitlab:viztor|@viztor]] was approved.
=== 2026-03-24 ===
* 23:18 [[gitlab:maryyann|@maryyann]] was approved.
* 23:01 [[gitlab:codenamenoreste|@codenamenoreste]] was approved.
* 13:36 [[gitlab:marc-maillard-wmse|@marc-maillard-wmse]] was approved.
* 07:39 "fred2675" was rejected (pending since 2025-12-23T07:39:11.380Z).
=== 2026-03-23 ===
* 14:51 [[gitlab:komla|@komla]] was approved.
* 05:51 "lunachuck43" was rejected (pending since 2025-12-22T05:50:17.862Z).
* 04:06 "reza110011" was rejected (pending since 2025-12-22T04:05:25.117Z).
=== 2026-03-20 ===
* 21:54 "mertgor" was rejected (pending since 2025-12-19T21:51:51.419Z).
* 20:57 "autanmahmah" was rejected (pending since 2025-12-19T20:54:51.678Z).
* 09:57 [[gitlab:nethahussain|@nethahussain]] was approved.
* 09:27 [[gitlab:piewriter|@piewriter]] was approved.
* 08:15 [[gitlab:dondersmooi|@dondersmooi]] was approved.
=== 2026-03-19 ===
* 21:03 "sayvhior" was rejected (pending since 2025-12-18T21:02:31.699Z).
=== 2026-03-18 ===
* 20:15 [[gitlab:martinmystere|@martinmystere]] was approved.
=== 2026-03-17 ===
* 02:51 "louperivois" was rejected (pending since 2025-12-16T02:50:48.197Z).
=== 2026-03-16 ===
* 12:54 "mokayaj857" was rejected (pending since 2025-12-15T12:53:39.015Z).
* 06:18 "roamer15" was rejected (pending since 2025-12-15T06:16:38.042Z).
=== 2026-03-14 ===
* 11:12 "umaramuhammad" was rejected (pending since 2025-12-13T11:10:44.004Z).
* 09:33 "akuma19" was rejected (pending since 2025-12-13T09:31:39.044Z).
* 07:06 [[gitlab:syunsyunminmin|@syunsyunminmin]] was approved.
=== 2026-03-12 ===
* 20:24 [[gitlab:11wb|@11wb]] was approved.
* 09:54 [[gitlab:bcxfu75k|@bcxfu75k]] was approved.
=== 2026-03-10 ===
* 09:12 [[gitlab:viktoriahillerudwmse|@viktoriahillerudwmse]] was approved.
=== 2026-03-06 ===
* 08:09 "vazhayilnewone" was rejected (pending since 2025-12-05T08:07:02.184Z).
=== 2026-03-04 ===
* 20:54 [[gitlab:elphie|@elphie]] was approved.
* 11:39 "ronaldahmed" was rejected (pending since 2025-12-03T11:37:47.492Z).
* 02:12 "ltslw" was rejected (pending since 2025-12-03T02:11:52.040Z).
=== 2026-03-02 ===
* 19:21 "dlopez350" was rejected (pending since 2025-12-01T19:20:38.918Z).
* 18:15 [[gitlab:lsandergreen|@lsandergreen]] was approved.
=== 2026-03-01 ===
* 10:51 [[gitlab:clintacc|@clintacc]] was approved.
=== 2026-02-28 ===
* 09:24 "cardboardlamp" was rejected (pending since 2025-11-29T09:22:03.947Z).
* 08:18 "wiki-pavan" was rejected (pending since 2025-11-29T08:16:24.184Z).
=== 2026-02-27 ===
* 20:45 "thisisrick25" was rejected (pending since 2025-11-28T20:42:24.454Z).
=== 2026-02-26 ===
* 13:57 "chuiimuiiofc" was rejected (pending since 2025-11-27T13:57:02.794Z).
* 13:54 "steffpro" was rejected (pending since 2025-11-27T13:52:10.859Z).
=== 2026-02-25 ===
* 21:24 "abubakarhabibudayyabu" was rejected (pending since 2025-11-26T21:22:37.776Z).
=== 2026-02-24 ===
* 05:00 "playboi" was rejected (pending since 2025-11-25T05:00:30.762Z).
=== 2026-02-23 ===
* 14:00 "alph65" was rejected (pending since 2025-11-24T13:59:00.797Z).
* 12:33 [[gitlab:robertsky|@robertsky]] was approved.
=== 2026-02-22 ===
* 00:30 "hp8p" was rejected (pending since 2025-11-23T00:29:24.741Z).
=== 2026-02-19 ===
* 16:45 "clayjar" was rejected (pending since 2025-11-20T16:44:48.380Z).
=== 2026-02-18 ===
* 22:18 "nexus" was rejected (pending since 2025-11-19T22:16:48.818Z).
* 12:00 "bernsteinnn" was rejected (pending since 2025-11-19T11:59:04.427Z).
=== 2026-02-17 ===
* 11:36 "jason2000-cpu" was rejected (pending since 2025-11-18T11:34:00.314Z).
=== 2026-02-16 ===
* 14:54 "smaurya" was rejected (pending since 2025-11-17T14:52:06.906Z).
=== 2026-02-15 ===
* 16:51 "kra-79" was rejected (pending since 2025-11-16T16:50:41.375Z).
=== 2026-02-14 ===
* 15:15 [[gitlab:mess|@mess]] was approved.
=== 2026-02-13 ===
* 13:57 "sopalsuemae957" was rejected (pending since 2025-11-14T13:55:16.921Z).
* 13:30 [[gitlab:wyslijp16-toolforge|@wyslijp16-toolforge]] was approved.
=== 2026-02-12 ===
* 16:30 "kristinagligoric" was rejected (pending since 2025-11-13T16:29:21.646Z).
* 03:33 [[gitlab:anyehansen|@anyehansen]] was approved.
* 02:21 [[gitlab:thejoyfultentmaker|@thejoyfultentmaker]] was approved.
=== 2026-02-10 ===
* 13:18 [[gitlab:db111|@db111]] was approved.
=== 2026-02-09 ===
* 19:06 "squirrel289" was rejected (pending since 2025-11-10T19:04:27.831Z).
=== 2026-02-06 ===
* 20:54 [[gitlab:gillux|@gillux]] was approved.
* 09:09 [[gitlab:lih|@lih]] was approved.
=== 2026-01-31 ===
* 16:21 [[gitlab:taxonbot1|@taxonbot1]] was approved.
=== 2026-01-28 ===
* 14:30 [[gitlab:ademola|@ademola]] was approved.
* 10:51 "watshell" was rejected (pending since 2025-10-29T10:51:01.521Z).
=== 2026-01-26 ===
* 23:06 "tavaresgmg" was rejected (pending since 2025-10-27T23:04:42.140Z).
=== 2026-01-25 ===
* 06:03 "cata" was rejected (pending since 2025-10-26T06:01:26.155Z).
=== 2026-01-24 ===
* 21:15 [[gitlab:wiegels|@wiegels]] was approved.
* 06:30 [[gitlab:blaquans|@blaquans]] was approved.
=== 2026-01-23 ===
* 16:27 [[gitlab:lerickson|@lerickson]] was approved.
* 10:15 "fran0035g" was rejected (pending since 2025-10-24T10:12:17.732Z).
=== 2026-01-22 ===
* 21:00 "hacksyn" was rejected (pending since 2025-10-23T20:59:15.982Z).
=== 2026-01-21 ===
* 17:30 [[gitlab:otcenas11|@otcenas11]] was approved.
=== 2026-01-19 ===
* 21:48 [[gitlab:amdrel|@amdrel]] was approved.
* 04:36 "rayalexa" was rejected (pending since 2025-10-20T04:35:02.094Z).
=== 2026-01-18 ===
* 15:45 "somya" was rejected (pending since 2025-10-19T15:43:43.701Z).
* 06:54 "sergg001" was rejected (pending since 2025-10-19T06:54:12.296Z).
=== 2026-01-16 ===
* 11:57 "zeejohsy" was rejected (pending since 2025-10-17T11:56:22.372Z).
* 04:45 "rocky25" was rejected (pending since 2025-10-17T04:43:33.180Z).
=== 2026-01-15 ===
* 16:39 "tiisu" was rejected (pending since 2025-10-16T16:37:18.438Z).
* 12:00 "noahalorwu" was rejected (pending since 2025-10-16T11:58:26.133Z).
* 10:39 "prjayaiuedu" was rejected (pending since 2025-10-16T10:37:16.947Z).
=== 2026-01-13 ===
* 17:21 [[gitlab:lwilson-ctr|@lwilson-ctr]] was approved.
=== 2026-01-12 ===
* 17:03 "stagietechs" was rejected (pending since 2025-10-13T17:02:25.281Z).
=== 2026-01-10 ===
* 19:06 "keerthisr" was rejected (pending since 2025-10-11T19:05:01.758Z).
=== 2026-01-09 ===
* 20:36 "lightb" was rejected (pending since 2025-10-10T20:34:20.264Z).
=== 2026-01-08 ===
* 19:42 [[gitlab:tbodt|@tbodt]] was approved.
* 13:57 [[gitlab:martynranyard|@martynranyard]] was approved.
=== 2026-01-07 ===
* 17:48 [[gitlab:santanuwiki25|@santanuwiki25]] was approved.
* 14:27 "dipanshu" was rejected (pending since 2025-10-08T14:26:10.794Z).
* 12:30 "adeolaadesina" was rejected (pending since 2025-10-08T12:29:49.592Z).
* 09:21 "tony-kamande" was rejected (pending since 2025-10-08T09:20:28.421Z).
* 06:18 "hninwuttyi" was rejected (pending since 2025-10-08T06:17:28.006Z).
* 05:09 "andume" was rejected (pending since 2025-10-08T05:07:18.582Z).
* 02:00 "mosope" was rejected (pending since 2025-10-08T01:59:54.800Z).
* 01:15 [[gitlab:tungstalite|@tungstalite]] was approved.
=== 2026-01-06 ===
* 18:24 "leerensucher" was rejected (pending since 2025-10-07T18:21:41.253Z).
* 14:54 "leonidlednev" was rejected (pending since 2025-10-07T14:53:07.273Z).
* 12:57 "alexandre-tingaud" was rejected (pending since 2025-10-07T12:54:27.206Z).
=== 2026-01-04 ===
* 21:33 [[gitlab:matr1x-101|@matr1x-101]] was approved.
* 15:18 "makjr" was rejected (pending since 2025-10-05T15:16:31.558Z).
* 14:09 "dakshq" was rejected (pending since 2025-10-05T14:08:40.608Z).
=== 2026-01-03 ===
* 20:42 [[gitlab:apehitkey|@apehitkey]] was approved.
* 18:00 [[gitlab:jeremyb|@jeremyb]] was approved.
* 14:09 [[gitlab:twelephant|@twelephant]] was approved.
=== 2026-01-01 ===
* 11:30 "shellstanislav" was rejected (pending since 2025-10-02T11:29:10.150Z).
=== 2025-12-30 ===
* 19:51 "camilojdiaz" was rejected (pending since 2025-09-30T19:49:24.913Z).
=== 2025-12-29 ===
* 16:03 "zied" was rejected (pending since 2025-09-29T16:01:30.415Z).
* 08:18 "rahulsidpradhan" was rejected (pending since 2025-09-29T08:17:02.849Z).
=== 2025-12-26 ===
* 09:48 "thembo42" was rejected (pending since 2025-09-26T09:45:15.033Z).
=== 2025-12-25 ===
* 14:03 "196936074751" was rejected (pending since 2025-09-25T14:02:31.367Z).
=== 2025-12-23 ===
* 16:21 "ngarnsworthy" was rejected (pending since 2025-09-23T16:20:41.211Z).
=== 2025-12-22 ===
* 12:39 "aza555" was rejected (pending since 2025-09-22T12:38:02.622Z).
=== 2025-12-20 ===
* 23:45 "saph" was rejected (pending since 2025-09-20T23:45:01.222Z).
=== 2025-12-19 ===
* 10:15 "vladdymoses" was rejected (pending since 2025-09-19T10:15:00.999Z).
* 07:15 "dirtylittlepoobah" was rejected (pending since 2025-09-19T07:13:55.537Z).
=== 2025-12-18 ===
* 16:24 [[gitlab:guyfawcus|@guyfawcus]] was approved.
=== 2025-12-17 ===
* 21:39 [[gitlab:holdyourhorses|@holdyourhorses]] was approved.
* 18:30 "prudencia" was rejected (pending since 2025-09-17T18:27:18.860Z).
* 02:24 "lottie" was rejected (pending since 2025-09-17T02:21:21.744Z).
=== 2025-12-16 ===
* 09:39 [[gitlab:melcatherine|@melcatherine]] was approved.
* 08:54 [[gitlab:leila237|@leila237]] was approved.
=== 2025-12-15 ===
* 18:27 [[gitlab:royalsailor|@royalsailor]] was approved.
* 09:39 [[gitlab:olaf8940|@olaf8940]] was approved.
* 09:39 "brianbybyby" was rejected (pending since 2025-09-15T09:37:45.430Z).
=== 2025-12-14 ===
* 20:21 [[gitlab:essa237|@essa237]] was approved.
* 16:42 [[gitlab:bovimacoco|@bovimacoco]] was approved.
=== 2025-12-13 ===
* 21:54 "mmns21" was rejected (pending since 2025-09-13T21:52:24.017Z).
* 20:33 "bugcrawler" was rejected (pending since 2025-09-13T20:31:09.211Z).
=== 2025-12-12 ===
* 14:39 "ruvchoudhary" was rejected (pending since 2025-09-12T14:36:16.167Z).
* 06:54 "rezadress" was rejected (pending since 2025-09-12T06:52:21.749Z).
=== 2025-12-10 ===
* 17:30 [[gitlab:itsmoon|@itsmoon]] was approved.
=== 2025-12-09 ===
* 15:42 [[gitlab:mercy-o|@mercy-o]] was approved.
=== 2025-12-06 ===
* 16:45 "jacquesradjabu" was rejected (pending since 2025-09-06T16:45:17.969Z).
* 11:27 [[gitlab:ikhitron|@ikhitron]] was approved.
=== 2025-12-01 ===
* 08:12 "halconmilenario21" was rejected (pending since 2025-09-01T08:12:10.262Z).
=== 2025-11-30 ===
* 21:06 [[gitlab:habs|@habs]] was approved.
=== 2025-11-29 ===
* 16:36 "bovimacoco" was rejected (pending since 2025-08-30T16:34:39.712Z).
* 00:45 [[gitlab:jjpmaster|@jjpmaster]] was approved.
=== 2025-11-24 ===
* 10:30 "alph65" was rejected (pending since 2025-08-25T10:28:40.957Z).
* 02:24 [[gitlab:yaron|@yaron]] was approved.
=== 2025-11-20 ===
* 16:06 "clayjar" was rejected (pending since 2025-08-21T16:04:54.450Z).
=== 2025-11-17 ===
* 21:09 [[gitlab:ankita97531|@ankita97531]] was approved.
=== 2025-11-16 ===
* 14:15 "commanderkefir" was rejected (pending since 2025-08-17T14:13:14.791Z).
* 08:21 "rehankhan78" was rejected (pending since 2025-08-17T08:19:44.896Z).
=== 2025-11-15 ===
* 14:36 "cyberscribe" was rejected (pending since 2025-08-16T14:34:27.230Z).
=== 2025-11-13 ===
* 04:21 "waddie96" was rejected (pending since 2025-08-14T04:19:27.461Z).
=== 2025-11-11 ===
* 06:42 [[gitlab:seanhoyland|@seanhoyland]] was approved.
=== 2025-11-10 ===
* 00:06 [[gitlab:jaredblumer|@jaredblumer]] was approved.
=== 2025-11-09 ===
* 22:36 "heinxiety" was rejected (pending since 2025-08-10T22:33:12.041Z).
=== 2025-11-07 ===
* 22:00 [[gitlab:forzagreen|@forzagreen]] was approved.
=== 2025-11-06 ===
* 16:57 [[gitlab:rsilvola|@rsilvola]] was approved.
=== 2025-11-04 ===
* 21:24 [[gitlab:devdoingdev|@devdoingdev]] was approved.
=== 2025-11-03 ===
* 17:48 "joewaleed98" was rejected (pending since 2025-08-04T17:46:12.191Z).
=== 2025-11-01 ===
* 18:00 "eliasempresas" was rejected (pending since 2025-08-02T17:58:04.412Z).
=== 2025-10-31 ===
* 18:51 [[gitlab:chaoticenby|@chaoticenby]] was approved.
* 04:33 "3ch310n" was rejected (pending since 2025-08-01T04:32:21.982Z).
=== 2025-10-30 ===
* 10:03 [[gitlab:tausheefhassan|@tausheefhassan]] was approved.
=== 2025-10-29 ===
* 14:54 "theap" was rejected (pending since 2025-07-30T14:52:12.066Z).
=== 2025-10-28 ===
* 06:06 [[gitlab:tanbiruzzaman|@tanbiruzzaman]] was approved.
=== 2025-10-27 ===
* 07:51 [[gitlab:jmoore111|@jmoore111]] was approved.
=== 2025-10-25 ===
* 21:09 [[gitlab:valor|@valor]] was approved.
* 21:03 [[gitlab:booksmurf|@booksmurf]] was approved.
* 02:48 "mystyc1" was rejected (pending since 2025-07-26T02:46:19.373Z).
=== 2025-10-24 ===
* 05:12 "aadarshmahesh" was rejected (pending since 2025-07-25T05:09:38.264Z).
=== 2025-10-22 ===
* 20:54 [[gitlab:janewanga|@janewanga]] was approved.
* 17:27 "abeljeevan" was rejected (pending since 2025-07-23T17:26:46.884Z).
* 16:12 "shrimpnaur" was rejected (pending since 2025-07-23T16:10:37.864Z).
=== 2025-10-21 ===
* 18:51 "jrmuizel" was rejected (pending since 2025-07-22T18:50:07.315Z).
* 09:33 [[gitlab:dpogorzelski|@dpogorzelski]] was approved.
=== 2025-10-17 ===
* 13:21 [[gitlab:blegodwin|@blegodwin]] was approved.
=== 2025-10-16 ===
* 14:51 [[gitlab:bahago|@bahago]] was approved.
* 14:12 "harikrishna0005" was rejected (pending since 2025-07-17T14:10:48.385Z).
* 14:09 "gauthammohanraj" was rejected (pending since 2025-07-17T14:08:47.643Z).
=== 2025-10-15 ===
* 13:48 [[gitlab:adwivedii|@adwivedii]] was approved.
* 13:18 [[gitlab:kimbrenekakande|@kimbrenekakande]] was approved.
* 13:03 "childmnajennifer" was rejected (pending since 2025-07-16T13:01:50.236Z).
* 05:06 "vssb4214" was rejected (pending since 2025-07-16T05:05:33.985Z).
=== 2025-10-14 ===
* 19:39 [[gitlab:afanyulionel|@afanyulionel]] was approved.
* 15:33 [[gitlab:sadrettin|@sadrettin]] was approved.
* 14:18 [[gitlab:tmwyk|@tmwyk]] was approved.
* 08:42 "yasu0796" was rejected (pending since 2025-07-15T08:41:26.453Z).
=== 2025-10-13 ===
* 16:09 [[gitlab:atlas0007|@atlas0007]] was approved.
=== 2025-10-11 ===
* 17:42 [[gitlab:techwizzie|@techwizzie]] was approved.
=== 2025-10-10 ===
* 19:03 [[gitlab:miiswom|@miiswom]] was approved.
* 16:06 [[gitlab:ninatakang|@ninatakang]] was approved.
=== 2025-10-09 ===
* 15:42 [[gitlab:jaykaneki|@jaykaneki]] was approved.
* 14:21 [[gitlab:lebogang|@lebogang]] was approved.
* 14:15 [[gitlab:kimondorose|@kimondorose]] was approved.
* 13:48 [[gitlab:joyakinyi|@joyakinyi]] was approved.
* 13:48 [[gitlab:dikshyashahi|@dikshyashahi]] was approved.
* 13:45 [[gitlab:obediobadiah|@obediobadiah]] was approved.
* 13:45 [[gitlab:system625|@system625]] was approved.
* 13:45 [[gitlab:rolalove|@rolalove]] was approved.
* 13:39 [[gitlab:olatundeawo|@olatundeawo]] was approved.
* 13:36 [[gitlab:danielchristlight|@danielchristlight]] was approved.
* 13:36 [[gitlab:dipanshu1223|@dipanshu1223]] was approved.
* 13:36 [[gitlab:aradhya|@aradhya]] was approved.
* 09:57 "bognd" was rejected (pending since 2025-07-10T09:55:48.661Z).
=== 2025-10-08 ===
* 23:36 [[gitlab:sopzy|@sopzy]] was approved.
* 23:03 [[gitlab:oluwatumininu|@oluwatumininu]] was approved.
* 19:39 [[gitlab:levon003|@levon003]] was approved.
* 15:24 [[gitlab:ritika-bhambri11|@ritika-bhambri11]] was approved.
* 13:45 [[gitlab:anbanguyen|@anbanguyen]] was approved.
* 13:36 [[gitlab:chumzine|@chumzine]] was approved.
* 13:27 [[gitlab:shr0x-ya|@shr0x-ya]] was approved.
* 12:45 [[gitlab:nurahwakili|@nurahwakili]] was approved.
* 03:42 "nazhiba" was rejected (pending since 2025-07-09T03:40:12.625Z).
* 02:12 "mafennel" was rejected (pending since 2025-07-09T02:11:40.598Z).
=== 2025-10-07 ===
* 22:54 [[gitlab:olusegunfaj|@olusegunfaj]] was approved.
* 21:30 [[gitlab:rona|@rona]] was approved.
* 21:09 [[gitlab:sandijigs|@sandijigs]] was approved.
* 13:36 "xisbajao" was rejected (pending since 2025-07-08T13:33:35.018Z).
* 01:36 "areczek94" was rejected (pending since 2025-07-08T01:35:40.633Z).
=== 2025-10-06 ===
* 19:21 "wmcarter2017" was rejected (pending since 2025-07-07T19:21:12.899Z).
=== 2025-10-05 ===
* 14:15 "meetmendapara" was rejected (pending since 2025-07-06T14:14:16.726Z).
=== 2025-10-04 ===
* 20:51 "nftbaee" was rejected (pending since 2025-07-05T20:50:57.688Z).
=== 2025-10-03 ===
* 06:12 [[gitlab:javiermonton|@javiermonton]] was approved.
=== 2025-10-02 ===
* 20:15 "talaqalotaibipmp" was rejected (pending since 2025-07-03T20:13:05.164Z).
=== 2025-10-01 ===
* 10:54 "bjensen" was rejected (pending since 2025-07-02T10:53:46.574Z).
* 02:45 "kowal1984" was rejected (pending since 2025-07-02T02:44:56.946Z).
=== 2025-09-30 ===
* 21:21 [[gitlab:kavaljeetsingh|@kavaljeetsingh]] was approved.
* 00:24 "adium" was rejected (pending since 2025-07-01T00:23:43.807Z).
=== 2025-09-28 ===
* 08:54 [[gitlab:pexerik|@pexerik]] was approved.
=== 2025-09-27 ===
* 13:57 [[gitlab:rubahhitamvukova|@rubahhitamvukova]] was approved.
=== 2025-09-26 ===
* 16:57 "algorithmic" was rejected (pending since 2025-06-27T16:56:17.480Z).
* 13:54 [[gitlab:shadabgdg|@shadabgdg]] was approved.
* 13:12 [[gitlab:spushpit|@spushpit]] was approved.
=== 2025-09-20 ===
* 14:06 "bwiki" was rejected (pending since 2025-06-21T13:59:14.749Z).
=== 2025-09-16 ===
* 05:39 [[gitlab:deepchirp|@deepchirp]] was approved.
=== 2025-09-15 ===
* 22:00 [[gitlab:noisk8|@noisk8]] was approved.
* 11:03 "ahonc" was rejected (pending since 2025-06-16T11:00:54.843Z).
=== 2025-09-13 ===
* 18:24 "a-ssh22" was rejected (pending since 2025-06-14T18:23:33.937Z).
* 12:36 [[gitlab:rajashreetalukdar|@rajashreetalukdar]] was approved.
* 00:45 [[gitlab:sumitsurai|@sumitsurai]] was approved.
=== 2025-09-12 ===
* 17:12 [[gitlab:suyash23|@suyash23]] was approved.
* 00:46 "remotetravel" was rejected (pending since 2025-06-13T00:44:08.171Z).
=== 2025-09-10 ===
* 21:09 "jancborchardt" was rejected (pending since 2025-06-11T21:06:30.759Z).
=== 2025-09-09 ===
* 17:03 [[gitlab:vwf|@vwf]] was approved.
* 06:36 [[gitlab:cactusisme|@cactusisme]] was approved.
=== 2025-09-08 ===
* 18:09 "birushandegeya" was rejected (pending since 2025-06-09T18:08:00.087Z).
* 16:27 "ngarnsworthy" was rejected (pending since 2025-06-09T16:24:37.213Z).
* 12:33 "zolgoyo" was rejected (pending since 2025-06-09T12:31:34.199Z).
=== 2025-09-06 ===
* 23:09 [[gitlab:jaishsingh913|@jaishsingh913]] was approved.
=== 2025-09-05 ===
* 21:45 [[gitlab:sakshi2|@sakshi2]] was approved.
* 20:42 "abdukhaliq1" was rejected (pending since 2025-06-06T20:40:42.023Z).
* 14:27 "beubsamy" was rejected (pending since 2025-06-06T14:27:06.781Z).
=== 2025-09-04 ===
* 23:27 "sdhehua" was rejected (pending since 2025-06-05T23:24:45.777Z).
* 19:00 [[gitlab:perry|@perry]] was approved.
* 11:24 "saintwolf" was rejected (pending since 2025-06-05T11:21:20.176Z).
=== 2025-09-02 ===
* 05:48 [[gitlab:aliu|@aliu]] was approved.
=== 2025-08-29 ===
* 13:30 "kksurendran066" was rejected (pending since 2025-05-30T13:27:48.755Z).
=== 2025-08-28 ===
* 22:18 "tauraamuix" was rejected (pending since 2025-05-29T22:16:08.228Z).
=== 2025-08-26 ===
* 19:03 [[gitlab:dikkulah|@dikkulah]] was approved.
=== 2025-08-22 ===
* 23:51 [[gitlab:khoroshun_mike|@khoroshun_mike]] was approved.
=== 2025-08-21 ===
* 07:39 [[gitlab:yuka|@yuka]] was approved.
=== 2025-08-19 ===
* 07:48 [[gitlab:zhaofjx|@zhaofjx]] was approved.
=== 2025-08-17 ===
* 14:27 "madhan13k" was rejected (pending since 2025-05-18T14:26:08.973Z).
=== 2025-08-15 ===
* 10:15 "mohammed_abukhadra" was rejected (pending since 2025-05-16T10:14:48.403Z).
=== 2025-08-11 ===
* 11:48 "hmmyesbro" was rejected (pending since 2025-05-12T11:45:24.350Z).
=== 2025-08-10 ===
* 13:15 [[gitlab:dactyl|@dactyl]] was approved.
=== 2025-08-09 ===
* 04:39 "xxxx100000" was rejected (pending since 2025-05-10T04:37:44.949Z).
=== 2025-08-08 ===
* 14:33 [[gitlab:josefanthony|@josefanthony]] was approved.
=== 2025-08-07 ===
* 23:42 [[gitlab:robins7|@robins7]] was approved.
* 21:42 [[gitlab:pols12|@pols12]] was approved.
* 17:15 "sbronson" was rejected (pending since 2025-05-08T17:15:08.834Z).
* 14:57 [[gitlab:alvindulle|@alvindulle]] was approved.
* 14:45 [[gitlab:xentos|@xentos]] was approved.
* 06:27 "jamesboste" was rejected (pending since 2025-05-08T06:25:14.793Z).
* 03:57 "ysun" was rejected (pending since 2025-05-08T03:55:07.348Z).
=== 2025-08-06 ===
* 21:51 "pols12" was rejected (pending since 2025-05-07T21:49:13.598Z).
* 01:51 "okeamah" was rejected (pending since 2025-05-07T01:48:50.114Z).
=== 2025-08-05 ===
* 09:15 "mobashir-2013" was rejected (pending since 2025-05-06T09:14:24.069Z).
=== 2025-08-01 ===
* 08:00 "douginamug" was rejected (pending since 2025-05-02T07:57:38.317Z).
=== 2025-07-31 ===
* 02:30 [[gitlab:ads|@ads]] was approved.
=== 2025-07-27 ===
* 13:15 "mrico2703" was rejected (pending since 2025-04-27T13:13:12.346Z).
* 10:17 [[gitlab:josephfrancis12|@josephfrancis12]] was approved.
* 10:17 [[gitlab:fuzzew|@fuzzew]] was approved.
* 05:57 [[gitlab:biscuitbobby|@biscuitbobby]] was approved.
* 05:48 [[gitlab:ecoholic|@ecoholic]] was approved.
=== 2025-07-26 ===
* 11:48 [[gitlab:chimnayyyy|@chimnayyyy]] was approved.
* 11:48 [[gitlab:alwinalbert|@alwinalbert]] was approved.
* 11:48 [[gitlab:hridyakk|@hridyakk]] was approved.
* 11:45 [[gitlab:gaurigupta21|@gaurigupta21]] was approved.
* 11:45 [[gitlab:binetaa|@binetaa]] was approved.
* 10:21 [[gitlab:jyothikat22|@jyothikat22]] was approved.
* 10:21 [[gitlab:zobotrombie|@zobotrombie]] was approved.
* 10:21 [[gitlab:flykrth|@flykrth]] was approved.
* 10:21 [[gitlab:mehrinshamim|@mehrinshamim]] was approved.
* 10:21 [[gitlab:aadhi13|@aadhi13]] was approved.
* 10:21 [[gitlab:malavikam05|@malavikam05]] was approved.
* 10:18 [[gitlab:nf609|@nf609]] was approved.
* 05:48 [[gitlab:nazalnihad|@nazalnihad]] was approved.
* 05:48 [[gitlab:naveen28204280|@naveen28204280]] was approved.
=== 2025-07-25 ===
* 09:49 [[gitlab:kasyap9|@kasyap9]] was approved.
* 09:30 [[gitlab:swayamagrahari|@swayamagrahari]] was approved.
=== 2025-07-24 ===
* 19:36 [[gitlab:madutgn|@madutgn]] was approved.
=== 2025-07-23 ===
* 20:09 [[gitlab:somerandomdeveloper|@somerandomdeveloper]] was approved.
=== 2025-07-22 ===
* 00:15 [[gitlab:iagoqnsi|@iagoqnsi]] was approved.
=== 2025-07-21 ===
* 17:30 [[gitlab:asadiqui|@asadiqui]] was approved.
* 16:39 [[gitlab:tryvix1509|@tryvix1509]] was approved.
* 04:27 [[gitlab:damian|@damian]] was approved.
=== 2025-07-20 ===
* 09:42 "mike-khoroshun" was rejected (pending since 2025-04-20T09:42:22.732Z).
=== 2025-07-17 ===
* 17:57 [[gitlab:haroldkrabs|@haroldkrabs]] was approved.
* 13:45 [[gitlab:envlh|@envlh]] was approved.
=== 2025-07-14 ===
* 10:24 [[gitlab:missguru|@missguru]] was approved.
* 00:57 "clarfonthey" was rejected (pending since 2025-04-14T00:56:32.626Z).
=== 2025-07-13 ===
* 01:01 [[gitlab:l235|@l235]] was approved.
=== 2025-07-11 ===
* 03:06 "rodavlas" was rejected (pending since 2025-04-11T03:05:45.590Z).
=== 2025-07-06 ===
* 00:09 "lakasa" was rejected (pending since 2025-04-06T00:06:28.469Z).
=== 2025-07-05 ===
* 21:54 "ctrlzvi" was rejected (pending since 2025-04-05T21:54:12.542Z).
* 14:30 "aminualiyu" was rejected (pending since 2025-04-05T14:27:22.617Z).
=== 2025-07-04 ===
* 03:15 [[gitlab:galstar|@galstar]] was approved.
=== 2025-07-02 ===
* 11:27 "vicolas11" was rejected (pending since 2025-04-02T11:25:12.682Z).
=== 2025-06-29 ===
* 23:12 "naomi723" was rejected (pending since 2025-03-30T23:09:24.630Z).
=== 2025-06-28 ===
* 16:21 "mudeh2372" was rejected (pending since 2025-03-29T16:18:27.057Z).
=== 2025-06-27 ===
* 23:18 "rony143" was rejected (pending since 2025-03-28T23:16:13.671Z).
* 22:21 [[gitlab:rluts|@rluts]] was approved.
=== 2025-06-26 ===
* 13:54 "creativegurus" was rejected (pending since 2025-03-27T13:52:41.706Z).
=== 2025-06-24 ===
* 17:42 [[gitlab:devjadiya|@devjadiya]] was approved.
* 14:00 "dominic-r" was rejected (pending since 2025-03-25T14:00:07.307Z).
=== 2025-06-21 ===
* 00:48 [[gitlab:vriaa|@vriaa]] was approved.
=== 2025-06-18 ===
* 15:21 "ayushkhati1" was rejected (pending since 2025-03-19T15:18:50.062Z).
=== 2025-06-17 ===
* 20:45 "chiomavero" was rejected (pending since 2025-03-18T20:44:13.967Z).
* 00:27 [[gitlab:eggroll97|@eggroll97]] was approved.
=== 2025-06-14 ===
* 20:57 "volvox" was rejected (pending since 2025-03-15T20:56:34.018Z).
=== 2025-06-13 ===
* 16:09 [[gitlab:supergrey|@supergrey]] was approved.
* 11:03 "chqaz" was rejected (pending since 2025-03-14T11:01:09.600Z).
* 10:24 [[gitlab:slong-wmf|@slong-wmf]] was approved.
* 10:15 "hearvox" was rejected (pending since 2025-03-14T10:13:13.112Z).
=== 2025-06-12 ===
* 15:18 "jlam" was rejected (pending since 2025-03-13T15:17:54.099Z).
=== 2025-06-09 ===
* 20:48 "dipanjansengupta" was rejected (pending since 2025-03-10T20:48:03.545Z).
* 19:27 [[gitlab:reggycelly|@reggycelly]] was approved.
* 14:51 "arendpieter" was rejected (pending since 2025-03-10T14:51:01.445Z).
* 13:21 [[gitlab:greenreaper|@greenreaper]] was approved.
* 09:33 [[gitlab:mmta|@mmta]] was approved.
* 08:03 "a-ssh22" was rejected (pending since 2025-03-10T08:03:08.111Z).
=== 2025-06-08 ===
* 21:06 "mm-episodenlistedlvaupdater" was rejected (pending since 2025-03-09T21:04:06.323Z).
=== 2025-06-06 ===
* 11:06 [[gitlab:olea|@olea]] was approved.
=== 2025-06-05 ===
* 20:33 [[gitlab:encodedwp|@encodedwp]] was approved.
* 15:00 [[gitlab:toluayo|@toluayo]] was approved.
* 13:51 [[gitlab:arnold_lup|@arnold_lup]] was approved.
* 11:54 "sdhehua" was rejected (pending since 2025-03-06T11:51:48.241Z).
=== 2025-06-03 ===
* 21:27 [[gitlab:wewakey|@wewakey]] was approved.
* 12:36 "hunsimon2" was rejected (pending since 2025-03-04T12:34:56.520Z).
* 11:54 "hunsimon" was rejected (pending since 2025-03-04T11:53:54.652Z).
=== 2025-06-02 ===
* 12:01 [[gitlab:jaimedes|@jaimedes]] was approved.
=== 2025-05-30 ===
* 18:00 "sathvik9105" was rejected (pending since 2025-02-28T17:59:42.867Z).
* 11:21 [[gitlab:tonythomas01|@tonythomas01]] was approved.
* 10:06 [[gitlab:gpsleo|@gpsleo]] was approved.
=== 2025-05-29 ===
* 22:12 [[gitlab:codynguyen1116|@codynguyen1116]] was approved.
=== 2025-05-28 ===
* 02:57 [[gitlab:saper|@saper]] was approved.
=== 2025-05-27 ===
* 21:06 [[gitlab:mohammed_qays|@mohammed_qays]] was approved.
* 15:33 "satanluimm" was rejected (pending since 2025-02-25T15:32:48.101Z).
=== 2025-05-26 ===
* 23:57 "seyedali220" was rejected (pending since 2025-02-24T23:56:17.621Z).
=== 2025-05-21 ===
* 11:12 [[gitlab:guilherme|@guilherme]] was approved.
=== 2025-05-19 ===
* 13:24 [[gitlab:emojiwiki|@emojiwiki]] was approved.
=== 2025-05-18 ===
* 00:00 "xidme" was rejected (pending since 2025-02-15T23:58:56.796Z).
=== 2025-05-17 ===
* 02:39 "kdh8219" was rejected (pending since 2025-02-15T02:36:32.237Z).
=== 2025-05-16 ===
* 15:09 [[gitlab:maxbinderwmf|@maxbinderwmf]] was approved.
=== 2025-05-15 ===
* 04:30 "inspectorzer0" was rejected (pending since 2025-02-13T04:27:33.179Z).
=== 2025-05-14 ===
* 17:42 [[gitlab:llugo|@llugo]] was approved.
=== 2025-05-13 ===
* 20:18 "mmta" was rejected (pending since 2025-02-11T20:17:23.407Z).
=== 2025-05-11 ===
* 20:51 "jad" was rejected (pending since 2025-02-09T20:49:07.333Z).
* 17:54 "nishchalsundan" was rejected (pending since 2025-02-09T17:52:25.761Z).
* 16:39 "mohammed_abukhadra" was rejected (pending since 2025-02-09T16:39:03.730Z).
=== 2025-05-09 ===
* 09:12 [[gitlab:sirchanmp|@sirchanmp]] was approved.
=== 2025-05-08 ===
* 08:18 [[gitlab:mengeditch|@mengeditch]] was approved.
=== 2025-05-07 ===
* 03:45 "xluffy" was rejected (pending since 2025-02-05T03:45:14.181Z).
=== 2025-05-06 ===
* 16:54 "punhaniabhishek" was rejected (pending since 2025-02-04T16:53:50.758Z).
* 09:36 [[gitlab:bmartinezcalvo|@bmartinezcalvo]] was approved.
=== 2025-05-02 ===
* 12:24 [[gitlab:tohaomg|@tohaomg]] was approved.
* 11:48 [[gitlab:mavrikant|@mavrikant]] was approved.
* 11:45 [[gitlab:daanvr|@daanvr]] was approved.
=== 2025-05-01 ===
* 09:09 "mjoerg" was rejected (pending since 2025-01-30T09:09:04.204Z).
=== 2025-04-30 ===
* 23:06 "sanskardubey" was rejected (pending since 2025-01-29T23:03:25.489Z).
=== 2025-04-29 ===
* 16:00 "geyslein" was rejected (pending since 2025-01-28T16:00:01.510Z).
=== 2025-04-26 ===
* 09:30 "anjali9027" was rejected (pending since 2025-01-25T09:28:07.064Z).
=== 2025-04-25 ===
* 18:00 "salahhazaa" was rejected (pending since 2025-01-24T17:58:30.030Z).
* 15:15 [[gitlab:yiming|@yiming]] was approved.
* 02:06 "mrchanmp" was rejected (pending since 2025-01-24T02:03:58.308Z).
=== 2025-04-23 ===
* 17:03 "rj2904" was rejected (pending since 2025-01-22T17:03:11.207Z).
* 14:21 "nischay33" was rejected (pending since 2025-01-22T14:19:21.081Z).
=== 2025-04-22 ===
* 19:27 "dj80" was rejected (pending since 2025-01-21T19:25:28.498Z).
* 14:30 [[gitlab:kaimamin|@kaimamin]] was approved.
* 09:57 "debo" was rejected (pending since 2025-01-21T09:54:47.955Z).
=== 2025-04-21 ===
* 12:24 "unshell" was rejected (pending since 2025-01-20T12:21:59.686Z).
=== 2025-04-18 ===
* 15:06 [[gitlab:spartanarbinger|@spartanarbinger]] was approved.
=== 2025-04-16 ===
* 03:09 "dewey" was rejected (pending since 2025-01-15T03:06:17.488Z).
=== 2025-04-15 ===
* 19:45 "emdadul" was rejected (pending since 2025-01-14T19:42:29.285Z).
=== 2025-04-14 ===
* 06:45 [[gitlab:bcampbell804|@bcampbell804]] was approved.
=== 2025-04-11 ===
* 06:27 [[gitlab:jvanderhoop|@jvanderhoop]] was approved.
=== 2025-04-10 ===
* 04:12 "bhai420" was rejected (pending since 2025-01-09T04:10:29.430Z).
=== 2025-04-09 ===
* 05:03 "austinvarshney" was rejected (pending since 2025-01-08T05:02:34.175Z).
=== 2025-04-06 ===
* 15:36 [[gitlab:elph|@elph]] was approved.
=== 2025-04-02 ===
* 10:33 [[gitlab:ozge|@ozge]] was approved.
=== 2025-03-31 ===
* 20:15 "demandkey" was rejected (pending since 2024-12-30T20:14:23.096Z).
* 15:18 [[gitlab:danyya|@danyya]] was approved.
=== 2025-03-28 ===
* 15:54 [[gitlab:rutsavi09|@rutsavi09]] was approved.
* 15:54 [[gitlab:ilanen1|@ilanen1]] was approved.
=== 2025-03-25 ===
* 19:27 [[gitlab:irfo|@irfo]] was approved.
* 11:54 [[gitlab:kmontalva-wmf|@kmontalva-wmf]] was approved.
* 04:33 [[gitlab:paul26|@paul26]] was approved.
* 04:18 "as1100k" was rejected (pending since 2024-12-24T04:18:06.813Z).
=== 2025-03-24 ===
* 11:33 "amzadkhankk" was rejected (pending since 2024-12-23T11:33:14.176Z).
=== 2025-03-23 ===
* 12:24 "wolfdo" was rejected (pending since 2024-12-22T12:23:35.056Z).
=== 2025-03-22 ===
* 09:45 [[gitlab:fjmustak|@fjmustak]] was approved.
=== 2025-03-20 ===
* 18:42 "sathishkokila" was rejected (pending since 2024-12-19T18:39:35.161Z).
* 17:03 [[gitlab:alien4444|@alien4444]] was approved.
* 15:27 [[gitlab:davidcoronel|@davidcoronel]] was approved.
=== 2025-03-19 ===
* 22:57 [[gitlab:r1f4t|@r1f4t]] was approved.
* 19:03 "daniel24ps" was rejected (pending since 2024-12-18T19:00:21.249Z).
* 14:18 [[gitlab:beepbooppenguin|@beepbooppenguin]] was approved.
=== 2025-03-18 ===
* 17:48 "rahulkundu1209" was rejected (pending since 2024-12-17T17:46:41.936Z).
* 08:15 "kirtisikka972" was rejected (pending since 2024-12-17T08:13:25.487Z).
=== 2025-03-15 ===
* 13:30 "tulspal_sidhu" was rejected (pending since 2024-12-14T13:29:10.606Z).
* 01:39 "peacedeadc" was rejected (pending since 2024-12-14T01:37:36.579Z).
=== 2025-03-14 ===
* 03:51 [[gitlab:chuckthebuck|@chuckthebuck]] was approved.
* 02:33 "yxngtrtxll" was rejected (pending since 2024-12-13T02:31:51.658Z).
=== 2025-03-13 ===
* 14:36 [[gitlab:iccander|@iccander]] was approved.
=== 2025-03-12 ===
* 23:21 "jokerchic36" was rejected (pending since 2024-12-11T23:21:00.670Z).
* 15:30 [[gitlab:naomi|@naomi]] was approved.
* 15:27 [[gitlab:cobi|@cobi]] was approved.
=== 2025-03-11 ===
* 12:42 "mohitvermaxx" was rejected (pending since 2024-12-10T12:40:56.967Z).
=== 2025-03-10 ===
* 16:51 [[gitlab:nanona15dobato|@nanona15dobato]] was approved.
=== 2025-03-09 ===
* 22:39 [[gitlab:jonkolbert|@jonkolbert]] was approved.
* 20:45 [[gitlab:urbanecmtest2|@urbanecmtest2]] was approved.
=== 2025-03-07 ===
* 16:54 [[gitlab:hswan|@hswan]] was approved.
* 14:42 [[gitlab:atitkov|@atitkov]] was approved.
* 00:42 [[gitlab:infrastruktur|@infrastruktur]] was approved.
=== 2025-03-06 ===
* 17:21 "johnmann" was rejected (pending since 2024-12-05T17:19:24.995Z).
=== 2025-03-05 ===
* 07:33 [[gitlab:monx9494|@monx9494]] was approved.
=== 2025-03-02 ===
* 21:21 "paul26" was rejected (pending since 2024-12-01T21:20:19.681Z).
=== 2025-03-01 ===
* 19:15 [[gitlab:izno|@izno]] was approved.
* 12:45 [[gitlab:nyerho|@nyerho]] was approved.
=== 2025-02-28 ===
* 18:27 [[gitlab:chuckonwumelu|@chuckonwumelu]] was approved.
* 13:09 "ashwinpraveengo" was rejected (pending since 2024-11-29T13:07:47.240Z).
* 00:18 "eduardoaugusto" was rejected (pending since 2024-11-29T00:17:43.372Z).
=== 2025-02-27 ===
* 20:39 "volkanurl" was rejected (pending since 2024-11-28T20:37:18.101Z).
=== 2025-02-24 ===
* 21:15 [[gitlab:feeglgeef|@feeglgeef]] was approved.
* 20:18 [[gitlab:piaanalysis2|@piaanalysis2]] was approved.
* 19:06 [[gitlab:dhardy|@dhardy]] was approved.
=== 2025-02-22 ===
* 19:27 [[gitlab:owuh|@owuh]] was approved.
=== 2025-02-19 ===
* 16:06 [[gitlab:artemkloko|@artemkloko]] was approved.
* 13:03 [[gitlab:jgafnea|@jgafnea]] was approved.
=== 2025-02-17 ===
* 16:33 [[gitlab:asmartkitten|@asmartkitten]] was approved.
=== 2025-02-16 ===
* 19:12 "gaurigupta21" was rejected (pending since 2024-11-17T19:11:07.416Z).
=== 2025-02-15 ===
* 01:18 [[gitlab:mediawiki-quickstart-ci|@mediawiki-quickstart-ci]] was approved.
=== 2025-02-14 ===
* 15:21 "nathanbnm" was rejected (pending since 2024-11-15T15:18:19.632Z).
=== 2025-02-13 ===
* 16:45 [[gitlab:priyanshuchahal|@priyanshuchahal]] was approved.
* 16:42 [[gitlab:ajhalili2006|@ajhalili2006]] was approved.
=== 2025-02-12 ===
* 23:21 "monkeypatch999" was rejected (pending since 2024-11-13T23:20:38.398Z).
* 06:36 [[gitlab:jainlakshita28|@jainlakshita28]] was approved.
=== 2025-02-11 ===
* 19:27 [[gitlab:matthewsm2|@matthewsm2]] was approved.
=== 2025-02-09 ===
* 16:15 "mohammed_abukhadra" was rejected (pending since 2024-11-10T16:15:18.361Z).
=== 2025-02-07 ===
* 21:33 "brennan" was rejected (pending since 2024-11-08T21:31:07.351Z).
=== 2025-02-06 ===
* 08:24 "mmta" was rejected (pending since 2024-11-07T08:22:36.724Z).
* 06:21 [[gitlab:bunnypranav|@bunnypranav]] was approved.
=== 2025-02-05 ===
* 22:39 "chrissteinchen" was rejected (pending since 2024-11-06T22:38:16.673Z).
=== 2025-02-03 ===
* 07:45 "edriiic" was rejected (pending since 2024-11-04T07:44:46.849Z).
* 01:12 "geppy" was rejected (pending since 2024-11-04T01:10:48.710Z).
=== 2025-02-02 ===
* 13:18 "funa-enpitu" was rejected (pending since 2024-11-03T13:15:46.065Z).
=== 2025-01-31 ===
* 23:42 "nfontes" was rejected (pending since 2024-11-01T23:39:41.755Z).
* 22:51 "sbronson" was rejected (pending since 2024-11-01T22:50:31.871Z).
* 00:42 [[gitlab:farid|@farid]] was approved.
=== 2025-01-27 ===
* 08:15 [[gitlab:eliza189|@eliza189]] was approved.
=== 2025-01-25 ===
* 09:51 [[gitlab:pamputt|@pamputt]] was approved.
=== 2025-01-23 ===
* 14:30 [[gitlab:lubianat|@lubianat]] was approved.
* 11:45 [[gitlab:bootsa|@bootsa]] was approved.
=== 2025-01-21 ===
* 05:09 "niko" was rejected (pending since 2024-07-21T16:10:01.377Z).
* 05:09 "thawizkid369777" was rejected (pending since 2024-07-18T17:42:44.493Z).
* 05:09 "sarthaksingh2" was rejected (pending since 2024-07-10T11:31:30.470Z).
* 05:09 "shriyakt" was rejected (pending since 2024-07-06T04:54:10.248Z).
* 05:09 "akshaya" was rejected (pending since 2024-07-06T04:04:51.488Z).
* 05:09 "alaka03aj" was rejected (pending since 2024-07-05T18:01:54.876Z).
* 05:09 "sulochanaviji-5049" was rejected (pending since 2024-07-01T05:58:00.427Z).
* 05:09 "nayanjnath" was rejected (pending since 2024-07-01T02:51:57.405Z).
* 05:09 "sd44" was rejected (pending since 2024-06-30T04:28:51.436Z).
* 05:09 "metavalent" was rejected (pending since 2024-06-29T01:37:14.210Z).
* 05:09 "wicloudx" was rejected (pending since 2024-06-28T11:51:23.335Z).
* 05:09 "debo" was rejected (pending since 2024-06-28T01:44:59.845Z).
* 05:09 "bwiki" was rejected (pending since 2024-06-23T14:15:38.032Z).
* 05:09 "toprak" was rejected (pending since 2024-06-23T11:35:50.819Z).
* 05:09 "iristeller" was rejected (pending since 2024-06-14T20:53:48.959Z).
* 05:09 "jcolvin" was rejected (pending since 2024-06-12T17:29:01.238Z).
* 05:09 "kalyan" was rejected (pending since 2024-06-07T07:52:46.993Z).
* 05:09 "bluecrystal" was rejected (pending since 2024-06-06T19:16:20.107Z).
* 05:09 "iftttrohit" was rejected (pending since 2024-06-04T12:08:50.818Z).
* 05:09 "pogpotato" was rejected (pending since 2024-06-03T17:58:21.684Z).
* 05:09 "cptlausebaer" was rejected (pending since 2024-05-31T18:53:27.692Z).
* 05:09 "hdevine825" was rejected (pending since 2024-05-31T17:04:18.279Z).
* 05:09 "anaghaa18" was rejected (pending since 2024-05-25T19:14:31.803Z).
* 05:09 "atharvanair04" was rejected (pending since 2024-05-25T14:24:52.825Z).
* 05:09 "anasvemmully" was rejected (pending since 2024-05-25T06:10:27.261Z).
* 05:09 "abhinavmohandas" was rejected (pending since 2024-05-25T06:05:24.825Z).
* 05:09 "kksurendran06" was rejected (pending since 2024-05-25T06:04:38.082Z).
* 05:09 "albertmarshall8896" was rejected (pending since 2024-05-23T09:32:05.462Z).
* 05:09 "akellison" was rejected (pending since 2024-05-17T02:07:24.229Z).
* 05:09 "mainowill" was rejected (pending since 2024-04-16T23:30:33.881Z).
* 05:09 "bzhqc" was rejected (pending since 2024-04-16T19:50:38.676Z).
* 05:09 "safan41" was rejected (pending since 2024-04-16T03:34:48.942Z).
* 05:09 "mgagat" was rejected (pending since 2024-04-16T03:21:51.764Z).
* 05:09 "okeamah" was rejected (pending since 2024-04-16T02:49:00.143Z).
* 05:09 "xuhao61" was rejected (pending since 2024-04-15T23:45:09.083Z).
* 04:47 "cybel" was rejected (pending since 2024-04-15T06:46:35.791Z).
=== 2025-01-20 ===
* 14:33 [[gitlab:your1|@your1]] was approved.
=== 2025-01-18 ===
* 10:09 [[gitlab:galrach600|@galrach600]] was approved.
* 02:51 [[gitlab:blankeclair|@blankeclair]] was approved.
=== 2025-01-17 ===
* 13:57 [[gitlab:dsantamaria|@dsantamaria]] was approved.
=== 2025-01-15 ===
* 17:12 [[gitlab:smartse|@smartse]] was approved.
=== 2025-01-14 ===
* 17:03 [[gitlab:naorleizer|@naorleizer]] was approved.
=== 2025-01-13 ===
* 02:45 [[gitlab:wolf20482|@wolf20482]] was approved.
=== 2025-01-12 ===
* 17:45 [[gitlab:tamzin|@tamzin]] was approved.
=== 2025-01-11 ===
* 15:24 [[gitlab:bargioni|@bargioni]] was approved.
* 14:30 [[gitlab:salelya|@salelya]] was approved.
* 10:15 [[gitlab:malakatshy|@malakatshy]] was approved.
* 05:21 [[gitlab:newmcpee|@newmcpee]] was approved.
=== 2025-01-09 ===
* 15:30 [[gitlab:gkyziridis|@gkyziridis]] was approved.
=== 2025-01-08 ===
* 16:21 [[gitlab:ukrface|@ukrface]] was approved.
=== 2024-12-28 ===
* 03:27 [[gitlab:twonum|@twonum]] was approved.
=== 2024-12-25 ===
* 06:09 [[gitlab:harsv567|@harsv567]] was approved.
=== 2024-12-21 ===
* 11:24 [[gitlab:amutha2002|@amutha2002]] was approved.
=== 2024-12-20 ===
* 19:51 [[gitlab:hridyeshgupta|@hridyeshgupta]] was approved.
* 10:00 [[gitlab:ro-shines|@ro-shines]] was approved.
* 08:09 [[gitlab:kesharwaniarpita|@kesharwaniarpita]] was approved.
=== 2024-12-18 ===
* 14:45 [[gitlab:soylacarli|@soylacarli]] was approved.
=== 2024-12-16 ===
* 20:33 [[gitlab:aleyasiddika1|@aleyasiddika1]] was approved.
=== 2024-12-15 ===
* 07:33 [[gitlab:abhishek02bhardwaj|@abhishek02bhardwaj]] was approved.
=== 2024-12-13 ===
* 13:18 [[gitlab:ashmitabathre204|@ashmitabathre204]] was approved.
=== 2024-12-10 ===
* 06:39 [[gitlab:ginaan|@ginaan]] was approved.
=== 2024-12-09 ===
* 05:45 [[gitlab:kallinavya|@kallinavya]] was approved.
* 00:54 [[gitlab:viserion-7|@viserion-7]] was approved.
=== 2024-12-08 ===
* 17:27 [[gitlab:wargo|@wargo]] was approved.
=== 2024-12-05 ===
* 11:15 [[gitlab:ranjithraj|@ranjithraj]] was approved.
=== 2024-12-02 ===
* 21:21 [[gitlab:a930913|@a930913]] was approved.
=== 2024-12-01 ===
* 02:39 [[gitlab:kingchristlike1|@kingchristlike1]] was approved.
=== 2024-11-21 ===
* 13:45 [[gitlab:sascha|@sascha]] was approved.
=== 2024-11-19 ===
* 16:36 [[gitlab:jly|@jly]] was approved.
=== 2024-11-15 ===
* 02:54 [[gitlab:danielyepezgarces|@danielyepezgarces]] was approved.
=== 2024-11-14 ===
* 14:15 [[gitlab:stimoroll|@stimoroll]] was approved.
=== 2024-11-09 ===
* 17:15 [[gitlab:f4udeveloper|@f4udeveloper]] was approved.
=== 2024-11-07 ===
* 19:15 [[gitlab:zulf|@zulf]] was approved.
* 05:33 [[gitlab:hassanamin|@hassanamin]] was approved.
=== 2024-11-06 ===
* 19:39 [[gitlab:daniuu|@daniuu]] was approved.
* 00:18 [[gitlab:rlopez-wmf|@rlopez-wmf]] was approved.
=== 2024-10-09 ===
* 14:45 [[gitlab:jtweed|@jtweed]] was approved.
* 10:24 [[gitlab:ifrahkh|@ifrahkh]] was approved.
* 09:06 [[gitlab:wikibayer|@wikibayer]] was approved.
=== 2024-10-06 ===
* 10:27 [[gitlab:keerthan16|@keerthan16]] was approved.
=== 2024-10-04 ===
* 07:45 [[gitlab:hakimi97|@hakimi97]] was approved.
=== 2024-09-30 ===
* 07:39 [[gitlab:ninjastrikers|@ninjastrikers]] was approved.
=== 2024-09-28 ===
* 17:30 [[gitlab:webrunner95|@webrunner95]] was approved.
=== 2024-09-18 ===
* 21:39 [[gitlab:elliottetzkorn|@elliottetzkorn]] was approved.
=== 2024-09-14 ===
* 22:06 [[gitlab:humptydumpty|@humptydumpty]] was approved.
=== 2024-09-06 ===
* 08:48 [[gitlab:mickabarber|@mickabarber]] was approved.
=== 2024-08-27 ===
* 17:36 [[gitlab:edgars|@edgars]] was approved.
=== 2024-08-22 ===
* 09:18 [[gitlab:antonkokhwmde|@antonkokhwmde]] was approved.
=== 2024-08-14 ===
* 19:21 [[gitlab:jfk|@jfk]] was approved.
=== 2024-08-13 ===
* 17:57 [[gitlab:daxserver|@daxserver]] was approved.
=== 2024-08-11 ===
* 09:57 [[gitlab:pauliesnug|@pauliesnug]] was approved.
=== 2024-08-10 ===
* 08:42 [[gitlab:ashig|@ashig]] was approved.
=== 2024-08-09 ===
* 14:09 [[gitlab:masssly|@masssly]] was approved.
=== 2024-08-05 ===
* 22:15 [[gitlab:mrtortue|@mrtortue]] was approved.
=== 2024-08-02 ===
* 16:21 [[gitlab:dsantini|@dsantini]] was approved.
=== 2024-07-31 ===
* 11:54 [[gitlab:cptviraj|@cptviraj]] was approved.
=== 2024-07-30 ===
* 19:09 [[gitlab:iniquity|@iniquity]] was approved.
* 10:00 [[gitlab:collins|@collins]] was approved.
=== 2024-07-27 ===
* 15:57 [[gitlab:songnguxyz|@songnguxyz]] was approved.
=== 2024-07-25 ===
* 12:36 [[gitlab:mszabo|@mszabo]] was approved.
* 09:21 [[gitlab:agarwalmahima|@agarwalmahima]] was approved.
=== 2024-07-24 ===
* 08:05 [[gitlab:dragoniez|@dragoniez]] was approved.
=== 2024-07-23 ===
* 06:54 [[gitlab:mirji|@mirji]] was approved.
=== 2024-07-16 ===
* 10:00 [[gitlab:lakejason0|@lakejason0]] was approved.
=== 2024-07-12 ===
* 11:33 [[gitlab:cn|@cn]] was approved.
* 08:12 [[gitlab:unchampignon|@unchampignon]] was approved.
=== 2024-07-07 ===
* 17:12 [[gitlab:agamyasamuel|@agamyasamuel]] was approved.
* 05:24 [[gitlab:kuldeepburjbhalaike|@kuldeepburjbhalaike]] was approved.
=== 2024-07-06 ===
* 11:18 [[gitlab:dibya|@dibya]] was approved.
* 04:54 [[gitlab:sarthakparashar|@sarthakparashar]] was approved.
=== 2024-07-05 ===
* 18:15 [[gitlab:vanshikarathi|@vanshikarathi]] was approved.
=== 2024-07-02 ===
* 19:00 [[gitlab:ebrahim|@ebrahim]] was approved.
=== 2024-07-01 ===
* 20:12 [[gitlab:rockingpenny4|@rockingpenny4]] was approved.
* 18:15 [[gitlab:balajijagadesh|@balajijagadesh]] was approved.
=== 2024-06-30 ===
* 18:24 [[gitlab:hrideshmg|@hrideshmg]] was approved.
* 07:18 [[gitlab:chanakyakumardas|@chanakyakumardas]] was approved.
* 06:30 [[gitlab:rihaan180|@rihaan180]] was approved.
=== 2024-06-27 ===
* 17:36 [[gitlab:driedmueller|@driedmueller]] was approved.
=== 2024-06-19 ===
* 12:57 [[gitlab:audreypenven|@audreypenven]] was approved.
=== 2024-06-16 ===
* 01:18 [[gitlab:roysmith|@roysmith]] was approved.
=== 2024-06-08 ===
* 02:45 [[gitlab:jleedev|@jleedev]] was approved.
=== 2024-06-03 ===
* 13:57 [[gitlab:afeder|@afeder]] was approved.
=== 2024-06-01 ===
* 10:54 [[gitlab:florianschmitt|@florianschmitt]] was approved.
=== 2024-05-30 ===
* 16:42 [[gitlab:krlsca|@krlsca]] was approved.
=== 2024-05-28 ===
* 11:24 [[gitlab:rickijay|@rickijay]] was approved.
=== 2024-05-26 ===
* 11:18 [[gitlab:ranjithsiji|@ranjithsiji]] was approved.
=== 2024-05-25 ===
* 07:24 [[gitlab:jony|@jony]] was approved.
=== 2024-05-23 ===
* 08:45 [[gitlab:lepticed7|@lepticed7]] was approved.
=== 2024-05-22 ===
* 20:42 [[gitlab:echecs|@echecs]] was approved.
=== 2024-05-21 ===
* 13:33 [[gitlab:mbs|@mbs]] was approved.
=== 2024-05-19 ===
* 18:06 [[gitlab:ionenlaser|@ionenlaser]] was approved.
=== 2024-05-18 ===
* 23:36 [[gitlab:mdaniels5757|@mdaniels5757]] was approved.
=== 2024-05-17 ===
* 08:54 [[gitlab:grapedog|@grapedog]] was approved.
=== 2024-05-08 ===
* 19:42 [[gitlab:kelhurd|@kelhurd]] was approved.
* 19:06 [[gitlab:khurd|@khurd]] was approved.
=== 2024-05-06 ===
* 19:48 [[gitlab:j3j5|@j3j5]] was approved.
* 12:06 [[gitlab:tk-999|@tk-999]] was approved.
=== 2024-05-05 ===
* 22:09 [[gitlab:pppery|@pppery]] was approved.
* 20:33 [[gitlab:sakretsu|@sakretsu]] was approved.
* 12:12 [[gitlab:waterquark|@waterquark]] was approved.
=== 2024-05-04 ===
* 09:03 [[gitlab:multichill|@multichill]] was approved.
* 07:42 [[gitlab:abaris|@abaris]] was approved.
=== 2024-05-03 ===
* 14:57 [[gitlab:maurusian|@maurusian]] was approved.
=== 2024-04-24 ===
* 05:48 [[gitlab:wolfinux|@wolfinux]] was approved.
=== 2024-04-23 ===
* 15:48 [[gitlab:dreamrimmer|@dreamrimmer]] was approved.
=== 2024-04-21 ===
* 06:51 [[gitlab:alon|@alon]] was approved.
=== 2024-04-17 ===
* 23:33 [[gitlab:derenrich|@derenrich]] was approved.
=== 2024-04-16 ===
* 17:18 [[gitlab:valcio|@valcio]] was approved.
=== 2024-04-14 ===
* 16:51 [[gitlab:wikilucas00|@wikilucas00]] was approved.
=== 2024-04-06 ===
* 12:48 [[gitlab:theprotonade|@theprotonade]] was approved.
=== 2024-04-02 ===
* 07:30 [[gitlab:bohuizhang|@bohuizhang]] was approved.
=== 2024-03-30 ===
* 13:36 [[gitlab:lpintscher|@lpintscher]] was approved.
=== 2024-03-26 ===
* 17:09 [[gitlab:eenabulele|@eenabulele]] was approved.
=== 2024-03-25 ===
* 14:27 [[gitlab:tuukka|@tuukka]] was approved.
=== 2024-03-24 ===
* 12:24 [[gitlab:firefly|@firefly]] was approved.
=== 2024-03-21 ===
* 19:33 [[gitlab:universal-omega|@universal-omega]] was approved.
=== 2024-03-17 ===
* 10:36 [[gitlab:bisel91|@bisel91]] was approved.
=== 2024-03-16 ===
* 10:09 [[gitlab:delord|@delord]] was approved.
* 00:42 [[gitlab:athulvis1|@athulvis1]] was approved.
=== 2024-03-15 ===
* 19:06 [[gitlab:ignaciorodrguez|@ignaciorodrguez]] was approved.
* 08:30 [[gitlab:peachey88|@peachey88]] was approved.
* 06:51 [[gitlab:derick|@derick]] was approved.
=== 2024-03-12 ===
* 15:06 [[gitlab:xiaoxiao|@xiaoxiao]] was approved.
=== 2024-03-06 ===
* 13:21 [[gitlab:desianabae1|@desianabae1]] was approved.
=== 2024-03-05 ===
* 19:21 [[gitlab:ep1c|@ep1c]] was approved.
* 16:33 [[gitlab:jasmine|@jasmine]] was approved.
=== 2024-03-02 ===
* 06:42 [[gitlab:potsdamlamb|@potsdamlamb]] was approved.
=== 2024-02-29 ===
* 23:18 [[gitlab:arandomname123|@arandomname123]] was approved.
* 18:03 [[gitlab:baba|@baba]] was approved.
* 17:48 [[gitlab:yfdyh000|@yfdyh000]] was approved.
* 03:09 [[gitlab:sds|@sds]] was approved.
=== 2024-02-27 ===
* 23:33 [[gitlab:lofhi|@lofhi]] was approved.
=== 2024-02-15 ===
* 19:45 [[gitlab:gergesshamon|@gergesshamon]] was approved.
=== 2024-02-14 ===
* 14:33 [[gitlab:philipnelson99|@philipnelson99]] was approved.
=== 2024-02-13 ===
* 13:06 [[gitlab:dringsim|@dringsim]] was approved.
=== 2024-02-12 ===
* 17:36 [[gitlab:haak|@haak]] was approved.
=== 2024-02-05 ===
* 17:33 [[gitlab:qwerfjkl|@qwerfjkl]] was approved.
* 17:14 [[gitlab:ahecht|@ahecht]] was approved.
=== 2024-02-01 ===
* 09:27 [[gitlab:arinaigum|@arinaigum]] was approved.
* 00:15 [[gitlab:jas42|@jas42]] was approved.
* 00:15 [[gitlab:edhu|@edhu]] was approved.
* 00:15 [[gitlab:marnanel|@marnanel]] was approved.
* 00:15 [[gitlab:ibrahemqasim|@ibrahemqasim]] was approved.
* 00:15 [[gitlab:amasotti|@amasotti]] was approved.
* 00:15 [[gitlab:deni|@deni]] was approved.
* 00:15 [[gitlab:cyber|@cyber]] was approved.
* 00:15 [[gitlab:saroj|@saroj]] was approved.
=== 2024-01-29 ===
* 21:42 [[gitlab:rgupta|@rgupta]] was approved.
=== 2024-01-07 ===
* 09:48 [[gitlab:lutrome|@lutrome]] was approved.
=== 2024-01-05 ===
* 20:48 [[gitlab:jinoytommanjaly|@jinoytommanjaly]] was approved.
* 02:51 [[gitlab:braunobruno|@braunobruno]] was approved.
* 01:08 [[gitlab:amorymeltzer|@amorymeltzer]] was approved.
* 01:08 [[gitlab:phi22ipus|@phi22ipus]] was approved.
=== 2024-01-03 ===
* 14:45 [[gitlab:gabina|@gabina]] was approved.
=== 2024-01-02 ===
* 13:18 [[gitlab:arthurtaylor|@arthurtaylor]] was approved.
=== 2023-12-23 ===
* 00:33 [[gitlab:aram|@aram]] was approved.
=== 2023-12-22 ===
* 16:24 [[gitlab:elpitareio|@elpitareio]] was approved.
=== 2023-12-21 ===
* 00:43 [[gitlab:bsadowski1|@bsadowski1]] was approved.
* 00:43 [[gitlab:ederporto|@ederporto]] was approved.
* 00:43 [[gitlab:sadraiiali|@sadraiiali]] was approved.
* 00:43 [[gitlab:wasp-outis|@wasp-outis]] was approved.
* 00:43 [[gitlab:bodhisattwa|@bodhisattwa]] was approved.
* 00:43 [[gitlab:air7538|@air7538]] was approved.
* 00:43 [[gitlab:anzx|@anzx]] was approved.
* 00:43 [[gitlab:tekask1903|@tekask1903]] was approved.
* 00:42 [[gitlab:kiwi-0x010c|@kiwi-0x010c]] was approved.
* 00:42 [[gitlab:mpaa|@mpaa]] was approved.
* 00:42 [[gitlab:kutay|@kutay]] was approved.
* 00:42 [[gitlab:wattmto|@wattmto]] was approved.
keunnezr0gz53gmy3hssdlnfx14s7gs
Wikidata Query Service/Migration/Rewrite of Label and Utility Services and Functions
0
460376
2445250
2444751
2026-08-09T03:48:17Z
Quiddity
1884
lang="sparql"
2445250
wikitext
text/x-wiki
This Wikitech page catalogs the rewrites needed to translate several utility query services provided by Blazegraph into portable, standards-conformant SPARQL 1.1. To port these queries, each Blazegraph-specific construct must be substituted with an equivalent SPARQL 1.1 pattern, restructured into a portable form, or (where no equivalent exists) removed with an explanatory note.
The rewrites described in this document cover six such constructs:
* The wikibase:label service
* Named subqueries
* Inline query hints
* bd:sample and bd:slice result-limiting predicates
* The wikibase:decodeURI function
* The wikibase:isSomeValue function
For each construct, the sections that follow explain its Blazegraph-specific semantics and specify the rewrite.
== wikibase:label ==
The wikibase:label service auto-resolves an entity's label, description, and aliases into result variables, honoring a language priority list with fallback. There are two modes used to define variable names - a manual and an automatic mode.
In the manual mode, the query writer declares the variable names. For example:
<syntaxhighlight lang="sparql">
SELECT ?item ?label ?description ?alternate WHERE {
?item wdt:P31 wd:Q515 .
SERVICE wikibase:label {
bd:serviceParam wikibase:language "de,en" .
?item rdfs:label ?label .
?item schema:description ?description .
?item skos:altLabel ?alternate . }
}
</syntaxhighlight>
This query produces results as shown below:
[[File:Rewrite Label-Fig 1.png|Figure 1]]
In automatic mode, no names are defined and the service projects specific names for any variables, ?x. The projected names are ?xLabel, ?xDescription, ?xAltLabel. A query could be written as:
<syntaxhighlight lang="sparql">
SELECT ?item ?itemLabel ?itemDescription ?itemAltLabel WHERE {
?item wdt:P31 wd:Q515 .
SERVICE wikibase:label {
bd:serviceParam wikibase:language "[AUTO_LANGUAGE],de,en". }
}
</syntaxhighlight>
This query produces the same results but with different names for the label, description and alias outputs.
When rewriting a query to produce equivalent results, the rewrite must consider:
* Language priority + fallback — wikibase:language is a comma-separated priority list where the first available language wins
** Usage of [AUTO_LANGUAGE] will be continued as a UI-only feature (it expands the list of languages to include the requester's UI/browser language)
* mul fallback — a label or alias under the multilingual code “mul” applies to all languages and is consulted after the listed languages
** There are no “mul” descriptions
* Allowance for unbound labels, descriptions and aliases
* QID fallback for labels (only)
** If no label exists in any requested language or mul, the entity ID (e.g., Q42) is returned
** There is no fallback for descriptions or altLabels (they are unbound if not found)
* skos:altLabel can be multivalued
** label service returns all aliases of the best language as one comma-joined string
* Applicable to both items and properties
At its most basic, the rewrite must drop the service call, determine each entity variable and its related label/description/alias variable names, and then return those variable values based on SPARQL OPTIONAL clauses.
Why SPARQL OPTIONALs? Because defining a raw triple pattern (such as ?item schema:description ?description) without OPTIONAL requires the item to have a description. If one is not defined, then the entity will be deleted from the result set. This is not the behavior of the wikibase:label service.
To rewrite the first query in this section, a naive approach (but one that illustrates what is happening “under the covers”) produces the following query:
<syntaxhighlight lang="sparql">
SELECT ?item ?label ?description ?alternate WHERE {
?item wdt:P31 wd:Q515 .
# --- label: priority de > en > mul, else QID ---
OPTIONAL { ?item rdfs:label ?label_de . FILTER(LANG(?label_de) = "de") }
OPTIONAL { ?item rdfs:label ?label_en . FILTER(LANG(?label_en) = "en") }
OPTIONAL { ?item rdfs:label ?label_mul . FILTER(LANG(?label_mul) = "mul") }
BIND(COALESCE(STR(?label_de), STR(?label_en), STR(?label_mul),
STRAFTER(STR(?item), "entity/")) AS ?label)
# --- description: priority de > en > mul, else unbound ---
OPTIONAL { ?item schema:description ?desc_de . FILTER(LANG(?desc_de) = "de") }
OPTIONAL { ?item schema:description ?desc_en . FILTER(LANG(?desc_en) = "en") }
BIND(COALESCE(STR(?desc_de), STR(?desc_en)) AS ?description)
# --- aliases: priority de > en, multi-valued → GROUP_CONCAT per language ---
OPTIONAL {
SELECT ?item (GROUP_CONCAT(DISTINCT ?a; SEPARATOR=", ") AS ?alt_de) WHERE {
?item skos:altLabel ?a . FILTER(LANG(?a) = "de")
} GROUP BY ?item
}
OPTIONAL {
SELECT ?item (GROUP_CONCAT(DISTINCT ?a; SEPARATOR=", ") AS ?alt_en) WHERE {
?item skos:altLabel ?a . FILTER(LANG(?a) = "en")
} GROUP BY ?item
}
OPTIONAL {
SELECT ?item (GROUP_CONCAT(DISTINCT ?a; SEPARATOR=", ") AS ?alt_mul) WHERE {
?item skos:altLabel ?a . FILTER(LANG(?a) = "mul")
} GROUP BY ?item
}
BIND(COALESCE(?alt_de, ?alt_en, ?alt_mul) AS ?alternate)
}
</syntaxhighlight>
There are several important things to note:
* All the resulting label/description labels are COALESCE’d as STR(?x) in order to remove the language tags
** This is needed since wikibase:label does not return language tags, but plain literals
** This is not needed for the aliases since GROUP_CONCAT returns a plain literal
* “mul” must be explicitly added as a language
* The COALESCE order is defined by the language priority with “mul” as the lowest priority
* The label’s COALESCE falls back to the entity QID if there are no labels in the specified languages
* Since aliases are multi-valued, each language’s results are GROUP_CONCAT'ed in their own GROUP BY ?item subqueries, then COALESCE'd for language priority
** There is a performance implication for this - the subqueries (SELECT clauses within the OPTIONALs) are always executed first and will concatenate alias names across all items with “de” or “en” aliases
*** This is because there are no binding constraints on ?item
** To improve the performance, the constraining triple(s) for ?item (in this case, <i>?item wdt:P31 wd:Q515</i>) should be moved inside the sub-SELECT scope
Executing the rewritten query obtains the following (equivalent but with different order) results:
[[File:Rewrite Label-Fig 2.png|Figure 2]]
The rewritten query was characterized as “naive” since it included the COALESCE. That forces an explicit ordering of the variable binding. However, it is not needed since the OPTIONAL clause is a left join. Once the first OPTIONAL binds ?label to the German label, the same variable in the next OPTIONAL's triple pattern is no longer a fresh binding. It becomes a join constraint: ?item rdfs:label ?label must now match the already-bound German literal, and FILTER(LANG(?label) = "en") on a @de literal is false. Therefore, the “en” OPTIONAL adds nothing and the German value survives. If a German label did not exist (?label is unbound), the English OPTIONAL is free to bind it; then mul; and so on. The left joins themselves implement "first match wins" which is what COALESCE was doing (making it redundant).
What remains is only the QID fallback for ?label and stripping the language tag. That can be written without COALESCE:
<syntaxhighlight lang="sparql">
BIND(IF(BOUND(?label), STR(?label), STRAFTER(STR(?item), "entity/")) AS ?itemLabel)
</syntaxhighlight>
Given this discussion, the rewritten query becomes:
<syntaxhighlight lang="sparql">
SELECT ?item ?label ?description ?alternate WHERE {
?item wdt:P31 wd:Q515 .
# label: de > en > mul into one shared ?lbl (left joins give priority), else QID
OPTIONAL { ?item rdfs:label ?lbl . FILTER(LANG(?lbl) = "de") }
OPTIONAL { ?item rdfs:label ?lbl . FILTER(LANG(?lbl) = "en") }
OPTIONAL { ?item rdfs:label ?lbl . FILTER(LANG(?lbl) = "mul") }
BIND(IF(BOUND(?lbl), STR(?lbl), STRAFTER(STR(?item), "entity/")) AS ?label)
# description: de > en into shared ?dsc, no mul and no fallback;
# STR(?dsc) to remove lang tag
OPTIONAL { ?item schema:description ?dsc . FILTER(LANG(?dsc) = "de") }
OPTIONAL { ?item schema:description ?dsc . FILTER(LANG(?dsc) = "en") }
BIND(STR(?dsc) AS ?description)
# aliases: de > en > mul into shared ?alternate, constrained, GROUP_CONCAT
OPTIONAL {
SELECT ?item (GROUP_CONCAT(DISTINCT ?a; SEPARATOR=", ") AS ?alternate)
WHERE { ?item wdt:P31 wd:Q515 . ?item skos:altLabel ?a . FILTER(LANG(?a) = "de")
} GROUP BY ?item }
OPTIONAL {
SELECT ?item (GROUP_CONCAT(DISTINCT ?a; SEPARATOR=", ") AS ?alternate)
WHERE { ?item wdt:P31 wd:Q515 . ?item skos:altLabel ?a . FILTER(LANG(?a) = "en")
} GROUP BY ?item }
OPTIONAL {
SELECT ?item (GROUP_CONCAT(DISTINCT ?a; SEPARATOR=", ") AS ?alternate)
WHERE { ?item wdt:P31 wd:Q515 . ?item skos:altLabel ?a . FILTER(LANG(?a) = "mul")
} GROUP BY ?item }
}
</syntaxhighlight>
The rewrite for the second query in this section (using automatic variable name binding and AUTO-LANGUAGE) uses the same overall pattern. But, there are a few items to note:
* The AUTO_LANGUAGE UI token will be maintained - so, it can be substituted as a language directly in the rewritten query (e.g., LANG(?x) = “AUTO_LANGUAGE”)
** Outside of the UI, there is no equivalent concept - all languages must be explicitly declared in the query
* If the query projected additional entity variables beyond ?item (e.g. ?country), each would need its own blocks for determining ?countryLabel, ?countryDescription and ?countryAlias
=== Label Service Results for an Unbound Variable ===
It is possible that the variable whose label is reported is not bound when the label results are returned. An example is a request for all plumbers' names and their mothers' names. Not all the plumbers' names are defined and many of their mothers are unknown.
For Blazegraph, this query would be written:
<syntaxhighlight lang="sparql">
SELECT ?a ?aLabel ?b ?bLabel WHERE {
?a wdt:P31 wd:Q5 .
?a wdt:P106 wd:Q252924 .
OPTIONAL { ?a wdt:P25 ?b }
SERVICE wikibase:label {
bd:serviceParam wikibase:language "en". }
}
</syntaxhighlight>
Using the rewrite pattern currently defined, the rewrite would appear as:
<syntaxhighlight lang="sparql">
SELECT ?a ?aLabel ?b ?bLabel WHERE {
?a wdt:P31 wd:Q5 .
?a wdt:P106 wd:Q252924 .
OPTIONAL { ?a wdt:P25 ?b }
OPTIONAL { ?a rdfs:label ?aLabel . FILTER ( lang(?aLabel) = 'en' ) }
OPTIONAL { ?b rdfs:label ?bLabel . FILTER ( lang(?bLabel) = 'en' ) }
}
</syntaxhighlight>
But, this query likely times out. Why? Each OPTIONAL clause is executed after "pushing down" all currently bound variables. The variable, ?a, is clearly bound by the triples in the second and third lines. So, the evaluation of the OPTIONAL clause for ?a's label is efficient. However, the variable, ?b, becomes bound in the <i>OPTIONAL { ?a wdt:P25 ?b }</i> clause, but that is not visible to the clause requesting ?b's label. The clause, <i>OPTIONAL { ?b rdfs:label ?bLabel . FILTER ( lang(?bLabel) = 'en' ) }</i>, becomes a request for all possible entities and their labels. Then, this is resolved using a left join across the OPTIONALs' results.
An approach to correct this is to check that the variable is bound before executing the OPTIONAL clause to retrieve its label. Taking this approach, the query would be rewritten as:
<syntaxhighlight lang="sparql">
SELECT ?a ?aLabel ?b ?bLabel WHERE {
?a wdt:P31 wd:Q5 .
?a wdt:P106 wd:Q252924 .
OPTIONAL { ?a wdt:P25 ?b }
BIND(IF(BOUND(?a), ?a, "unbound") AS ?al)
OPTIONAL { ?al rdfs:label ?aLabel . FILTER ( lang(?aLabel) = 'en' ) }
BIND(IF(BOUND(?b), ?b, "unbound") AS ?bl)
OPTIONAL { ?bl rdfs:label ?bLabel . FILTER ( lang(?bLabel) = 'en' ) }
}
</syntaxhighlight>
This works as desired, returning a set of results equivalent to that from Blazegraph.
Clearly, there are multiple ways to modify the query above.
# Since ?a is clearly bound, the IF(BOUND(?a), ...) check is not needed
# Or, the first OPTIONAL clause could be rewritten as:
::<syntaxhighlight lang="sparql">
OPTIONAL { ?a wdt:P25 ?b .
OPTIONAL { ?b rdfs:label ?bLabel . FILTER ( lang(?bLabel) = 'en' ) } }
</syntaxhighlight>
However, a straightforward and simple BIND(IF(BOUND(?x), ... check requires less analysis and is still performant.
=== Query Rewrite Pattern for wikibase:label ===
Collecting all the requirements and considerations above, the rewrite pattern can be explained as:
* Determine the languages to be output, splitting wikibase:language values using a comma delimiter
** And their priority order
* Determine the names of the projected variables
** Which then defines the types of OPTIONAL “blocks” needed - label/description/alias
* Create those blocks based on the pattern shown in the last query above (BIND checking for a bound variable and then the OPTIONAL clauses requesting the label in the specified and prioritized languages), and substitute them in the query in place of the wikibase:label service
** For the <i>alias</i> labels, taking care to re-insert the constraining triple(s) for the variable(s) in the OPTIONAL clauses' SELECT subqueries' body
*** This is needed since those SELECT subqueries are executed first and their results used in the OPTIONAL clauses left-join operation at the end of processing
== Named Subqueries ==
Although SPARQL supports nested subqueries, it does not support assigning them a name or specifically “including” them at specific (sometimes multiple) points in a scoping query. Blazegraph allows this using the keywords, INCLUDE, AS and WITH. For example:
<syntaxhighlight lang="sparql">
SELECT ?person ?employer
WITH {
SELECT ?person WHERE {
?person wdt:P31 wd:Q5 ; wdt:P106 wd:Q901 . # humans who are scientists
}
} AS %scientists
WHERE {
INCLUDE %scientists .
?person wdt:P108 ?employer .
}
</syntaxhighlight>
The query first executes the <i>WITH { … } AS %name</i> subquery, and does this only once. Its results are materialized into a named temporary set, and every <i>INCLUDE %name</i> joins against that same cached set. It's effectively a SQL Common Table Expression (CTE) using WITH … AS.
Named subqueries serve three purposes:
* Clarity - cleanly defining the query
* Reuse - allowing an expensive subquery to be executed once and used in several places
* Join-order control - forcing the subquery to run first, rather than letting the query optimizer to insert it
Unfortunately, SPARQL 1.1 has no equivalent to named subqueries. The portable equivalent is an inline sub-SELECT that duplicates the code of the WITH clause and replaces each <i>INCLUDE %name</i> with the body of the sub-SELECT. For the example above:
<syntaxhighlight lang="sparql">
SELECT ?person ?employer WHERE {
{ SELECT ?person WHERE {
?person wdt:P31 wd:Q5 ; wdt:P106 wd:Q901 .
} }
?person wdt:P108 ?employer .
}
</syntaxhighlight>
The most important question is whether the subqueries are simply copied wherever they are referenced. And, for ordinary, deterministic queries, the answer is yes. The sub-SELECT body can be pasted at each INCLUDE site<ref>This is actually a worst case rewrite since the pattern can be simplified to only paste within each unique query scope. This is discussed in more detail in the Query Rewrite sub-section.</ref>.
Again, unfortunately, "simply copied" hides three problems:
* Complexity, when a sub-query is copied multiple times and the query text “explodes”
* Performance, since the compute-once guarantee is lost
** Blazegraph evaluates the block once regardless of the number of references to it
** Instead, if it is inlined at N sites, then the query is evaluated N times
** For a “cheap” subquery, this is mainly irrelevant
** For an “expensive” subquery referenced several times, a large performance cost could be incurred
* Correctness, when the sub-query is non-deterministic or mints blank nodes AND occurs multiple times
** Non-deterministic queries include functions such as RAND() and UUID()/STRUUID()
*** When Blazegraph executes the sub-query once, every INCLUDE sees the same bound variables and values
*** Instead, if inlined multiple times, each copy re-evaluates those functions and returns different values
** The sub-query may return variables as blank nodes
*** As above, when Blazegraph executes the sub-query once, every INCLUDE sees the same bound variables and bnode values
*** Instead, if inlined multiple times, each copy returns different bnode identities
** Other examples of non-deterministic results are subqueries that use:
*** A LIMIT but no ORDER BY, making the result selection defined at execution time
*** GROUP_CONCAT, since the SPARQL aggregate functions operate over unordered sets
*** SAMPLE()
*** A federated service call to an external endpoint
**** If there are N separate calls to the federated service, results could vary between the calls due to changes in the data, a LIMIT on the federated result, or other non-deterministic behavior at the endpoint
*** Blazegraph-specific sampling utilities such as bd:slice and bd:sample
If non-deterministic or blank node results are included in the solution set, <i>AND the named subquery is used more than once</i>, it cannot just be copied but must first be executed/materialized. Then, its results can be manually added back into the query as a VALUES block.
For example, rewriting the scientist query above using the VALUES approach (even though this is deterministic), results in the following:
<syntaxhighlight lang="sparql">
SELECT ?person ?employer WHERE {
VALUES ?person { wd:Q4118502 wd:Q100156581 … } # Results of running %scientists once
?person wdt:P108 ?employer .
}
</syntaxhighlight>
There are a few obvious problems with the VALUES approach. First, it is a two-step manual or programmatic process. Second, it is cumbersome when the result set is large (thousands of results).
=== Query Rewrite Pattern for Named Subqueries ===
For rewriting a named subquery, its contents (within the WITH { … } AS clause) are replaced by an inlined <i>{ SELECT variables WHERE { … } }</i> clause. If there is only a single reference to the sub-query, the result of inlining is equivalent to the original query.
However, questions arise when there are multiple references to the sub-query and these occur in multiple scopes<ref>Examples of different scopes in a SPARQL query are having nested sub-queries or separate UNION branches.</ref>. One solution to this is to rewrite the sub-query in each scope. Further complications then arise only when the results of the sub-query are non-deterministic and this is visible in the externalized / observable result variables.
There is complexity in determining these two characteristics (non-deterministic and projected). To determine them, the query’s algebra tree<ref>Obtained using Jena ARQ or rdflib libraries</ref> must be analyzed. This can be used to first define whether multiple scopes are involved. If the enclosing graph pattern for each reference to a named subquery is the same - then there is only one scope (and no rewrite problems). However, if two INCLUDEs have different enclosing patterns, then their potential results need to be examined.
A simple way to determine if the results are non-deterministic due to generated blank nodes is to actually run the sub-query and see if any results are reported with type = “bnode”<ref>It is important to distinguish Wikidata skolemized blank nodes, and query engine-minted nodes. Wikidata’s blank nodes will be returned in a query binding as type="uri" with a specific, consistent IRI value.</ref>. If so and the subquery is referenced in multiple scopes, then that query is not eligible for rewriting.
Other simple checks for non-deterministic results is whether the subquery includes bd:slice or bd:sample (which have approximate rewrites), or a LIMIT without an ORDER BY.
Lastly, if the subquery uses SAMPLE(), RAND(), UUID(), STRUUID() or GROUP_CONCAT(), and if that use occurs in multiple scopes, and if the results of the function calls are observable in the results, then the query is <b>ineligible</b> for rewrite. Being “observable” occurs if the functions appear in the SELECT clause, affect which result sets are returned (used in FILTER/HAVING or ORDER BY under a LIMIT/OFFSET), or feed a projected BIND.
Examples of these are:
* A SELECT variable (for example, SELECT ?k (SAMPLE(?v) AS ?s))
* A FILTER/HAVING clause (e.g., HAVING(SAMPLE(?v) > 5))
* An ORDER BY (SAMPLE(?v)) ... LIMIT n clause (since the results which “clear” the LIMIT are dependent on the SAMPLE selection)
=== What about QLever’s Materialized Views? ===
QLever has defined a feature (still in beta at the time of writing) called [https://docs.qlever.dev/materialized-views/ materialized views] which seems similar to Blazegraph’s named subqueries. Both provide the capability to create a named, compute-once result set that is referenced by name. However, there are several differences and explicit issues. The most important of these are:
* SPARQL 1.1 compliance:
** Neither Blazegraph named subqueries nor QLever materialized views are SPARQL 1.1 compliant
** Also, the creation of a materialized view requires server-level access (not allowed on WDQS)
* Definition location:
** A named subquery (WITH … AS %x) is defined inside the query itself - it is self-contained and recomputed every run
** A materialized view is created out of band and stored on disk (dependent on server-side state)
* Lifetime and freshness:
** Since named subqueries are recomputed on every query execution, they are never stale
** Materialized views are persistent and must be recreated if they go stale
== Query Hints ==
QLever does not expose a Blazegraph-style query-hint system and this is not a part of the SPARQL 1.1 standard. So, the rewrite pattern for Blazegraph hint:* directives is to strip them entirely.
It is important to note that QLever's optimizer can be influenced by the structure of a SPARQL query. It is recommended to use subqueries to force early evaluation of a portion of a query. When enclosing a block inside a SPARQL subquery (SELECT ... WHERE { ... }), QLever will execute that block before joining its results with the remainder of the query.
== bd:sample ==
bd:sample is a Blazegraph service that selects a bounded subset of solutions from a graph pattern <i>without</i> fully materializing and sorting the entire set of results. Its purpose is to keep otherwise-expensive queries under the WDQS timeout by working on a sample instead of the full data.
For example, the following query returns 2000 “locations” (<i>wdt:P276 ?value</i>) and orders them by the number of distinct ?items that reference them. (Note however, that this is the ordering within the sample set of results and not across Wikidata.)
<syntaxhighlight lang="sparql">
SELECT ?value (COUNT(DISTINCT ?item) as ?count) WHERE {
SERVICE bd:sample {
?item wdt:P276 ?value .
bd:serviceParam bd:sample.limit 2000 . }
} GROUP BY ?value
ORDER BY DESC(COUNT(DISTINCT ?item))
</syntaxhighlight>
The query pattern defined within the service binds the relevant variables (in this case, ?item and ?value). These can then be used in the surrounding SPARQL.
The possible service parameters are:
{| class="wikitable"
|-
! Parameter !! Usage !! Default
|-
| bd:sample.limit || Maximum number of result sets to return || 100
|-
| bd:sample.sampleType || “RANDOM”, “EVEN” (evenly spaced sample), “DENSE” (first N) || RANDOM
|-
| bd:sample.seed || RANDOM's seed || 0
|}
Rewriting the service depends on the sampleType parameter, but none of the rewrites preserve the performance characteristics of bd:sample. For all rewrites, the engine must evaluate the inner pattern first and run to completion.
The closest SPARQL 1.1 query to that shown above is:
<syntaxhighlight lang="sparql">
SELECT ?value (COUNT(DISTINCT ?item) AS ?count) WHERE {
{ # Translate the query pattern within the service to a SELECT query
SELECT ?item ?value WHERE {
?item wdt:P276 ?value .
} ORDER BY RAND() # Default is RANDOM, so apply that ordering
LIMIT 2000 # LIMIT 2000
}
} GROUP BY ?value
ORDER BY DESC(COUNT(DISTINCT ?item))
</syntaxhighlight>
The above query does complete on QLever, but a query that returns a larger result set (for example, asking for a subset of humans, Q5) or is more complex could exceed the 60 second timeout. In all cases, applying a LIMIT to the inner query is needed to force a cut-off. That is also why bd:sample has a limit parameter.
It is valuable to note that the SPARQL 1.1 query could be changed to remove the ORDER BY RAND() statement and just use LIMIT. That would definitely improve the performance, and would mimic choosing the “DENSE” sampleType of bd:sample.
“EVEN” is the last sampleType value that could be used. To replicate the behavior requires positional/windowing logic that standard SPARQL 1.1 cannot express. However, the result can be approximated with the RANDOM rewrite above (a uniform random sample is usually an acceptable stand-in for an evenly-spaced one).
There is another service parameter that we did not discuss - seed. This would be used for reproducibility. Here is a Blazegraph query using that parameter:
<syntaxhighlight lang="sparql">
SELECT ?item WHERE {
SERVICE bd:sample {
?item wdt:P31 wd:Q5 . # ?item is an instance of human
bd:serviceParam bd:sample.limit 1000 .
bd:serviceParam bd:sample.sampleType "RANDOM" .
bd:serviceParam bd:sample.seed 42 . # fixed seed → same 1000 every run
}
}
</syntaxhighlight>
The rewrite is:
<syntaxhighlight lang="sparql">
SELECT ?item WHERE {
{ SELECT ?item WHERE {
?item wdt:P31 wd:Q5 .
} ORDER BY MD5(CONCAT("42", STR(?item))) # "42" plays the role of bd:sample.seed
LIMIT 1000
}
}
</syntaxhighlight>
Note that the above query can successfully execute on a QLever server in under 60 seconds. So, simply having a large intermediate result set does not necessarily equate to query timeout.
=== Query Rewrite Pattern for bd:sample ===
These steps summarize the rewrite patterns for bd:sample:
* Lift the query <pattern> out of SERVICE bd:sample {…} and rewrite it as an inner SELECT … WHERE { <pattern> }, projecting the variables present in the outer query
* Append LIMIT N (the bd:sample.limit or use the default, 100, if it was omitted)
* Add an ORDER BY if the sampleType is “RANDOM” or “EVEN”
** RAND() if a seed is not specified
** MD5(CONCAT(seed, STR(?key))) otherwise
* Leave the surrounding query (GROUP BY, COUNT, etc.) unchanged
== bd:slice ==
bd:slice is a Blazegraph service that provides efficient, stable pagination over a single triple pattern. It walks through a very large result set in deterministic chunks without re-sorting on each request. Its purpose is to keep otherwise-expensive queries under the WDQS timeout by working on a subset instead of the full data.
The most “mechanically” equivalent SPARQL 1.1 pattern is to perform the query with an ORDER BY, and OFFSET / LIMIT. However, this becomes increasingly non-performant as the OFFSET size grows - since OFFSET skips rows in the result set by first generating and then discarding them.
Another way to rewrite the query using SPARQL 1.1 is to create an “index” and use that to segment the results. Taking this approach each page carries a similar performance cost.
For example, the following Blazegraph query is used to return result sets of 10000 humans. Subsequent calls increase the offset by 10000 and add 10000 to the limit.
<syntaxhighlight lang="sparql">
SELECT ?item WHERE {
SERVICE bd:slice {
?item wdt:P31 wd:Q5 . # single triple pattern (required)
bd:serviceParam bd:slice.offset 0 .
bd:serviceParam bd:slice.limit 10000 .
}
}
</syntaxhighlight>
To rewrite this query using SPARQL 1.1 and a query-relevant index requires some manual construction … The largest value of the value from a run is inserted into a FILTER on the index.
For the query above, the first iteration is written as:
<syntaxhighlight lang="sparql">
SELECT ?item WHERE {
?item wdt:P31 wd:Q5 .
} ORDER BY (str(?item))
LIMIT 10000
</syntaxhighlight>
The second iteration becomes:
<syntaxhighlight lang="sparql">
SELECT ?item WHERE {
?item wdt:P31 wd:Q5 .
FILTER(STR(?item) > "http://www.wikidata.org/entity/Q100310806")
} ORDER BY STR(?item)
LIMIT 10000
</syntaxhighlight>
Where the entity, Q100310806, is the “last” ?item output by the previous iteration.
It is important to note that because of the indexes and performance characteristics of QLever, it is also possible to execute the query unconstrained by order or limits.
<syntaxhighlight lang="sparql">
SELECT ?item WHERE {
?item wdt:P31 wd:Q5 .
}
</syntaxhighlight>
A total of 13M+ results are output, and indicates that bd:slice was a performance optimization required by Blazegraph as opposed to addressing a query-related need.
As for bd:sample, there is another service parameter for bd:slice that should be discussed - range. In range mode, bd:slice doesn't return rows of results, but returns a single value - a count of the solutions that is computed from index boundaries. A variable is named and bd:slice.range estimates a count. Note that word, estimates. This is problematic since it is usually an overestimate and done as another performance optimization.
Consider these two Blazegraph queries:
<syntaxhighlight lang="sparql">
SELECT ?range WHERE {
SERVICE bd:slice {
?item wdt:P569 ?dob . # single triple pattern
bd:serviceParam bd:slice.range ?range . # ?range ← fast index count of matches
}
}
</syntaxhighlight>
And:
<syntaxhighlight lang="sparql">
SELECT ?range WHERE {
SERVICE bd:slice {
?item wdt:P569 ?dob .
FILTER(?dob >= "1900-01-01"^^xsd:dateTime && ?dob < "2000-01-01"^^xsd:dateTime)
bd:serviceParam bd:slice.range ?range .
}
}
</syntaxhighlight>
Both of these queries return a ?range of approximately 8074600 (varying +/- 3).
However, a corresponding query using SPARQL 1.1, gives exact results.
<syntaxhighlight lang="sparql">
SELECT (COUNT(*) AS ?range) WHERE {
?item wdt:P569 ?dob .
FILTER(?dob >= "1900-01-01T00:00:00"^^xsd:dateTime &&
?dob < "2000-01-01T00:00:00"^^xsd:dateTime)
}
</syntaxhighlight>
Which returns a count of 5457644. With the FILTER removed, the count is 8074572 (pretty close to the overall Blazegraph estimate!).
=== Query Rewrite Pattern for bd:slice ===
The following list defines the rewrite patterns for bd:slice with limit and offset, or with range:
* For limit/offset:
** Evaluate first if limit/offset are actually needed. For many queries, QLever’s indexes and performance are sufficient to remove the need for this optimization.
** If the query does timeout, then pagination is needed. The simplest portable form is to use ORDER BY <key> + OFFSET + LIMIT, bumping OFFSET with each query request. But, note that OFFSET cost grows with depth.
*** For deep paging, it is recommended to use keyset pagination. Pick a deterministic sort key (e.g. STR(?subject)), ORDER BY it, issue an initial LIMIT query, then for each subsequent page add <i>FILTER(STR(?subject) > “last_value_from_previous_query”)</i> with the same ORDER BY and LIMIT.
*** Keyset replaces the growing OFFSET with an index seek
*** Remember - do not use both OFFSET and the cursor FILTER together
** This is accomplished as a manual rewrite based on the bd:slice query pattern with the addition of a cursor/index (such as STR(?subject_var)) and ORDER BY statement using this index. A manual evaluation of the results is required - to determine the largest index value returned for the query. After the first query, each one has to include a FILTER() statement based on the largest index value returned, along with a LIMIT and OFFSET.
* For range:
** Use the query pattern as-is and <i>SELECT (COUNT(*) AS ?range)</i>
== wikibase:decodeURI() ==
wikibase:decodeUri reverses percent-encoding (for example, é → %C3%A9, comma → %2C, etc.) in a string. It is the explicit inverse of SPARQL 1.1's ENCODE_FOR_URI and turns IRIs such as Wikipedia/Commons sitelink URLs into human-readable strings.
Basically, this is a function that is used in BINDs and projected variable clauses to create human output. An exemplary Blazegraph query that translates a Russian Wikipedia sitelink to readable characters is:
<syntaxhighlight lang="sparql">
SELECT ?item ?sitelink ?decodedURI ?title WHERE {
?sitelink schema:about ?item ;
schema:isPartOf <https://ru.wikipedia.org/> .
BIND(wikibase:decodeUri(STR(?sitelink)) AS ?decodedURI)
} LIMIT 10
</syntaxhighlight>
The following illustrates one of the results:
* ?item: http://www.wikidata.org/entity/Q14760226
* ?sitelink: https://ru.wikipedia.org/wiki/%D0%90%D0%B1%D1%83%D0%BB%D1%8C_%D0%A5%D0%B0%D0%B4%D0%B6%D0%B4%D0%B6%D0%B0%D0%B4%D0%B6_I
* ?decodedURI: https://ru.wikipedia.org/wiki/Абуль_Хаджджадж_I
Regrettably, there is no equivalent decodeURI functionality in SPARQL 1.1. It cannot be created since that would require iterating over a string to assemble <i>multi-byte UTF-8 sequences</i> (as seen above, %C3%A9 becomes é). SPARQL does not support loops or byte-level operations, so a complete and correct decoder is not possible.
The only approach is to translate the URI after the SPARQL results are returned. Every programming language has a URL decode component. Using one of these languages (such as python or perl) on a Mac or Linux system, the command line instruction is:
:<syntaxhighlight lang="text">
python3 -c "import sys,urllib.parse as u; print(u.unquote(sys.argv[1]))" "uri_string_to_decode"
</syntaxhighlight>
OR
:<syntaxhighlight lang="text">
perl -CS -MURI::Escape -MEncode -e 'print decode_utf8(uri_unescape($ARGV[0])),"\n"' "uri_string_to_decode"
</syntaxhighlight>
On Windows, it can be accomplished in Powershell by:
:<syntaxhighlight lang="shell-session">
pwsh -c "[Uri]::UnescapeDataString('uri_string_to_decode')"
</syntaxhighlight>
An alternate strategy when decoding Wikipedia sitelinks is to instead retrieve the “title” from the value of the predicate, schema:name. For example:
<syntaxhighlight lang="sparql">
SELECT ?item ?title WHERE {
?article schema:about ?item ;
schema:isPartOf <https://ru.wikipedia.org/> ;
schema:name ?title .
} LIMIT 10
</syntaxhighlight>
Note that this does not return the full URI, but only the last path segment. So, at best, it is an approximation to aid in human readability.
=== Query Rewrite Pattern for wikibase:decodeURI ===
There is no general rewrite pattern. decodeURI() should be handled manually, at the client, using the command line instructions shown above.
Specifically for sitelinks (identified as the subject variable of a triple with the predicate, schema:about), a decoded subset of the full URI could be obtained by retrieving the value of the variable’s schema:name predicate.
<i>?something schema:about ?item … BIND(wikibase:decodeUri(STR(?something)) AS ?decodedURI)</i> becomes <i>?something schema:about ?item ; schema:name ?decodedTitle</i>. However, further string manipulations of ?decodedURI will require updating since ?decodedURI and ?decodedTitle are not equivalent.
== wikibase:isSomeValue() ==
wikibase:isSomeValue detects and filters statements that have an "unknown value" or "some value" (specifically, their snak type is 'somevalue') rather than a specific, concrete data item. This is used when data is known to exist, but the exact value is unspecified or unknown.
For example, the following query searches for all persons where their cause of death is specifically defined as "unknown".
<syntaxhighlight lang="sparql">
SELECT ?person ?deathStatement WHERE {
?person wdt:P31 wd:Q5 . # Find instances of humans
?person p:P509 ?deathStatement . # Get the cause of death stmt
?deathStatement ps:P509 ?causeValue . # Extract the value
# Keep only the rows where the cause of death is "unknown"
FILTER (wikibase:isSomeValue(?causeValue))
}
</syntaxhighlight>
The query returns 235 results. Obviously, there should be many more results if one searched on all persons who are dead but have no cause of death (<i>?person wdt:P570 ?dateOfDeath . FILTER NOT EXISTS {?person wdt:P509 ?cause}</i>). But this is a different query than asking where cause of death is specifically set to "unknown".
The query above is rewritten using the skolemized IRI for "someValue". That IRI has the format:
* Namespace: <nowiki>http://www.wikidata.org/.well-known/genid/</nowiki>
* Suffix: An MD5_HASH (e.g., d50f382e536becbbbc0987cd91bf1287) which is defined so that every "Some Value" assertion is distinct from other blank nodes or unknown values
Given this specific format, the query above can be rewritten as:
<syntaxhighlight lang="sparql">
SELECT ?person ?personLabel ?deathStatement WHERE {
?person wdt:P31 wd:Q5 . # Find instances of humans
?person p:P509 ?deathStatement . # Get the cause of death stmt
?deathStatement ps:P509 ?causeValue . # Extract the value
# Keep only the rows where the cause of death is an "unknown"
FILTER (CONTAINS(STR(?causeValue), "wikidata.org/.well-known/genid"))
}
</syntaxhighlight>
Which also returns 235 results.
=== Query Rewrite Pattern for wikibase:isSomeValue ===
The rewrite pattern is straightforward. A reference to wikibase:isSomeValue(?xxx) is translated to (CONTAINS(STR(?xxx), "wikidata.org/.well-known/genid")). This works whether the function is used in a BIND, IF or FILTER statement or as a projected variable.
A rewrite of a reference to wikibase:isSomeValue(?xxx, ?yyy) is translated to individual rewrites of each argument, joined with ||, and wrapped in parens so precedence stays correct (for example, under an enclosing operator like negation).
== Footnotes ==
[[Category:WDQS]]
nn9y8vij4xalkvc0oy4b98n8vgoshug